---
title: "SafeCode Consulting - Some Assembly Required : Filling an RTOS Gap"
description: "Porting a Linux driver to an RTOS without spinlocks came down to two assembly instructions wrapped in C. Why C alone can't do it, and how the pair works."
url: "https://safecodetech.com/insights/articles/design-2-lil1-some-assembly-required.html?tmpl=component"
date: "2026-10-06T20:47:49-04:00"
language: "en-US"
---

## Breadcrumbs

[Home](https://safecodetech.com/) > Insights > [Articles](https://safecodetech.com/insights/articles.html) > Some Assembly Required : Filling an RTOS Gap

## Some Assembly Required : Filling an RTOS Gap

Details By Max Hinkley Max Hinkley October 06, 2026

##### *Part 1 of 3 of the Locks-in-Layers series*

Not so very long ago, I was faced with an interesting problem. It happens now and then. In this case, I was porting a Linux Ethernet device driver to an RTOS. The driver was sizeable and complex. It transferred large blocks of data using DMA. There was a schedule crunch (rare, I know), and a previous attempt at this same task had failed miserably. I had to get a stable port as quickly as possible.

The driver used several Linux mechanisms that our RTOS lacked. Among others, it relied on spinlocks and tasklets. The Linux version of the driver was mature and stable. Rather than restructuring an unfamiliar codebase around the primitives we did have available (events, mutexes, and semaphores), I decided that it would be much faster and safer (with respect to avoiding defects) to simply implement the capabilities that our RTOS had omitted. The strategy worked, and I completed the task faster than most thought possible.

As I worked through the driver, a dependency chain became clear. The mechanisms I needed to replicate (data-structure locks, tasklets, timed functions) could all be implemented using a spinlock.

### What is a spinlock?

A spinlock is a lightweight software mechanism that prevents multiple processes from simultaneously executing within the same area of code, or any other code protected by the same lock. In its simplest application, it works much like a critical section, but since a single “lock” can be used to guard multiple sections, it has advantages. For example, it can be used to guard several short segments in a longer workflow to prevent uncoordinated changes of state. This makes it a very useful tool in multicore processing, which is the problem domain that it was created for.

The name “spinlock” was given because the process attempting to enter the guarded area “spins”; that is, it doesn’t sleep or wait – it just keeps retrying until it gets through. This is because the lock is intended to be held for very short periods, typically only a few machine instructions. In real-world concepts, this reminds me of some high-security entry gates at DoD facilities where one person can enter (acquiring the lock), the security guard checks ID and inspects carry-ins (a few simple operations), then the person exits (releasing the lock), enabling the next person to enter.

In software, in order to prevent interference among the contenders for the lock, there needs to be a “flag” that indicates whether the lock is being held. If two processes read the flag, nearly simultaneously, and it shows that the lock is not being held, they may both believe that they have the right of entry, and both set the flag. That doesn’t work. What is needed here is a guarantee that only one can gain entry. That requires the ability to read, modify, and write the flag in a single, atomic operation. Opcodes that support atomic operations are available on most modern MPUs but they are not normally made accessible by any clear API.

### Where C falls short

A large part of C's appeal for systems work is that unoptimized compilation is close to a direct assembly-language translation: that is, you get most of the speed of hand-written assembly without having to know the intimate details of every MPU you are programming. That translation only covers constructs that are built into the language. Nothing in C's grammar expresses an atomic read-modify-write. The “stdatomic.h” library header, made available in C11, isn’t universally supported, and certainly wasn’t available on my project. Some RTOSes expose atomic operations, but few provide an ISR-safe API.

So how can one write C code to build a spinlock when the critical piece is not available natively in C? Inline assembly provides the way. Let’s be clear: assembly language is not something to be used casually in a C codebase. If applied judiciously, it can provide immense power without obscuring the logic of the code. It can be used to simply expand C’s vocabulary. More on this later.

### The Atomic Solution

Most processors do not offer a general atomic read-modify-write that can be applied to arbitrary operations. Instead, many provide a pair of instructions that can be used to make almost any simple operation work as though it were atomic. They emulate atomicity.

The x86 architecture does not follow the pair pattern. Instead, its `CMPXCHG` is truly atomic in a single instruction: compare a location with an expected value and, if equal, swap in a new one. The operation is powerful and can also be used for spinlocks; in fact, I will be emulating that instruction in a function called replace_if_equal, which is the basis of my spinlock implementation. That is all that will be said about this processor family.

As stated, on many architectures the missing word is a pair of instructions rather than one. Here are some common examples:

ARM:`ldrex r0, [r1] ; load, and reserve the address`
`strex r2, r0, [r1] ; store only if the reservation holds`
`; r2 = 0 on success, nonzero on failure`

RISC-V:`lr.w t0, (a0) ; load-reserved`
`sc.w t1, t2, (a0) ; store-conditional; t1 = 0 on success`

PowerPC:`lwarx r5, 0, r3 ; load word, and reserve the address`
`stwcx. r6, 0, r3 ; store only if the reservation holds`

### The Instruction Pair, and How It Emulates Atomicity

The common theme for each of these instruction pairs is how it works. Different architectures may differ on the exact semantics and what comprises a memory region, but overall, they are very similar.

The first instruction performs a load from the memory location, and simultaneously raises a “reservation” flag associated with the memory region of that address. Rather than target-specific opcodes, I’ll generically refer to that operation as load_and_reserve. Once the reservation is placed, any modification of content within that memory region will “break the reservation” (reset the flag).

The second instruction writes a value to a memory location only if that region has a standing reservation, and provides a boolean flag to indicate whether the write succeeded. I’ll call that one store_if_reserved.

With two simple instructions, one can retrieve a value, operate on it, then attempt to store the result with an indication of whether the sequence was interrupted. It isn’t exactly atomic, but the feedback mechanism allows you to try again or take a different path if it fails.

### The Essential Blend

Now we know that these instructions are available via assembly language. What shall we do with that information?

In books, examples of these mechanisms are typically presented in pure assembly language. For my purposes, I wanted to avoid embedding long sections of assembly language code.

Another common approach that I’ve seen is to insert the inline assembly statements at each point where they are used. This switching of languages can also be difficult to follow.

It is often hard to gain acceptance of assembly language in a C codebase. Many C developers were exposed to some arbitrary assembly language in school, but for most, it is something they have never actually used in their professional lives. Reviewing an algorithm written in an assembly language you aren’t familiar with can be like reading it in a foreign language that cannot be understood without a language dictionary, and even then, full understanding can be elusive.

Since C is fully capable of expressing most of the logic very efficiently, it is much better to write in the language most accessible to the team, and to create a C layer only for those concepts that do not exist natively within the language.

Another way to view this is that we, as humans, can hear a parable from another language or culture. While it may seem to make sense, we can miss the point of the underlying lesson. Then, when we are given context and background about a specific cultural reference, the whole thing comes together, and we fully understand. That’s the aim here: to provide C with some added context and vocabulary, that we may gain a better understanding. To do this, I expose the instruction pair as a pair of inline C functions, each wrapping one of the instructions, while providing them with meaningful names. This way, the algorithms built on these primitives are not just technological magic. They are applicable, real-world concepts.

### Expanding the C vocabulary

The names for the wrappers were already provided above. What that leaves is just the implementations. Here, I’ll cover just the ARM example, using standard GCC inline assembly syntax:

`typedef unsigned int value_t;`
`typedef value_t reserved_t;`

`static inline value_t load_and_reserve( volatile reserved_t* res ) {`
`value_t v;`
`__asm__ volatile ( "ldrex %0, [%1]"`
`: "=r"(v)`
`: "r"(res)`
`: "memory" );`
`return v;`
`}`

`static inline bool store_if_reserved( volatile reserved_t* res, value_t val ) {`
`value_t failed;`
`__asm__ volatile ( "strex %0, %2, [%1]"`
`: "=&r"(failed)`
`: "r"(res), "r"(val)`
`: "memory" );`
`return ( failed == 0 );`
`}`

With those two short functions, we now have everything needed to create pseudo-atomic operations completely in the native C language.

### Next Time Around

The next installment, [Atomic-Powered Functions from Two Tiny Primitives](https://safecodetech.com/insights/articles/insights/articles/design-3-lil2-atomic-powered-functions), looks at the simple atomic operations that can be built on these basic functions. After that, we’ll see the solution I arrived at for my project.
