---
title: "SafeCode Consulting - Locks in Layers: Different Strokes"
description: "Porting a Linux driver to an RTOS without spinlocks came down to two assembly instructions wrapped in C. Why C alone can't do it, and how the pair works."
url: "https://safecodetech.com/insights/articles/design-4-lil3-locks-in-layers.html?tmpl=component"
date: "2026-10-06T20:47:56-04:00"
language: "en-US"
---

## Breadcrumbs

[Home](https://safecodetech.com/) > Insights > [Articles](https://safecodetech.com/insights/articles.html) > Locks in Layers: Different Strokes

## Locks in Layers: Different Strokes

Details By Max Hinkley Max Hinkley October 06, 2026

##### *Part 3 of 3 in the Locks-in-Layers series*

Now that we've covered the fundamental principles, it's time to show how the locks are created.

In part 1, [Some Assembly Required: Filling an RTOS Gap](https://safecodetech.com/insights/articles/insights/articles/design-2-lil1-some-assembly-required), I described how my approach to porting a driver from Linux to an RTOS was to replicate some of the functionality of Linux on the RTOS, specifically spinlocks, locks, & tasklets. That may seem like a huge undertaking, but it was accomplished pretty quickly. That article goes on to explain how it all rested on two assembly language primitives that were then wrapped up into C functions.

In part 2, [Atomic-Powered Functions from Two Tiny Primitives,](https://safecodetech.com/insights/articles/insights/articles/design-3-lil2-atomic-powered-functions) I covered how those two functions could be used to create “atomic” operations in C, and how those operations in turn enabled more complex mechanisms.

In this final installment, I’ll explain how I built the spinlocks and locks that I needed from the primitives. Tasklets and timed functions were a bit more complex and needed some other foundations, as well. If there is interest, maybe that will come in a follow-up article.

### More on Spinlocks

As it was explained in the earlier article, a spinlock is a guard used around one or more short segments of code to prevent multiple processes from executing that code simultaneously. All processes awaiting entry through the spinlock sit and “spin” as they wait; they don’t suspend or sleep, so the guarded segments must be kept very short to prevent prolonged blocking of those other processes. In single-core systems, spinlocks should generally be shared only among peer processes, as a higher-priority process (or ISR) “spinning”, may result in deadlock if the design isn’t carefully considered. An ISR that won’t be pre-empted can simply query whether the lock is held by another process, and choose its action from there.

Here is the spinlock in a nutshell:

`SPINLOCK`

`The lock is one shared value: FREE or HELD.`

`To acquire the lock:`
`Repeat:`
`Look at the lock.`
`If it is FREE:`
`Try to change it from FREE to HELD in a single atomic step.`
`If the change succeeded:`
`You own the lock. Stop repeating.`
`If it failed:`
`Someone else got there first. Go around again.`

`If it is HELD:`
`Go around again.`

`Once you own the lock, perform the guarded steps as quickly as possible.`

`To release the lock:`
`Simply set the value back to FREE.`

### Layer One: The Primitive Pair

This has been covered in depth in the previous articles, so we won’t dwell on it here.

`/* This gets the current value and places a reservation on the memory region (often a full page) */`
`static inline value_t load_and_reserve( volatile reserved_t* res ) {`
`value_t v;`

`/* ASSEMBLY_LANGUAGE_LOAD_AND_RESERVE(res, v);` */

`return v;`
`}`

`/* This stores a value in res ONLY if the region is reserved */`
`static inline bool store_if_reserved( volatile reserved_t* res, value_t val ) {`
`value_t result;`

`/* ASSEMBLY_LANGUAGE_STORE_IF_RESERVED(res, val, result);` */

`return ( result == STORE_SUCCEEDED );`
`}`

There is one other piece that is useful here. It is technically not assembly language, but the inline assembler provides the access. There are C rules about how compilers may reorder variable accesses. Volatile accesses may never be reordered among themselves, but the compiler may optimize by reordering non-volatile accesses, and this means that they can be moved outside of the locks we build, completely defeating the purpose of the lock. The compiler fails to see the reason for the sequence. There is a simple fix:

`/* This creates a compiler barrier (GCC / Clang syntax). It emits no instruction and has no run-time cost;`
`* it only stops the compiler from moving memory accesses across this point. Volatile accesses are never`
`* reordered relative to each other, but compilers are permitted to move accesses to non-volatile data`
`* across them when they see no dependency. That defeats the purpose of a lock, and this function`
`* serves as a localized preventative.`
`* @note This does not order accesses as seen by other processors (multicore).`
`*/`
`static inline void suppress_reordering( void ) { __asm__ volatile("" ::: "memory"); }`

The code listings in this article do not make use of `suppress_reordering`; however, the source code download at the end of the article puts it to use in all the right places.

### Layer Two: The atomic compare and exchange

This is the key piece, central to the function of the lock as a guardian. The name is different (replace_if_match), and we use the simplified replace_if_zero wrapper.

`/**`
`* Replace the value of res with val, only if *res matches the value of cmp.`
`* @param res [IN/OUT] The data location`
`* @param cmp [IN] The value to compare with the current content.`
`* @param val [IN] The value to be conditionally written.`
`* @param act [OUT] The actual value of res at exit.`
`* @return true iff *res contained cmp, and val was written; otherwise false.`
`*/`
`bool replace_if_match( volatile reserved_t* res, const value_t cmp, const value_t val, value_t* act ) {`

`value_t v;`

`do {`
`v = load_and_reserve( res );`

`if ( v != cmp ) {`
`*act = v;`

`return false;`
`}`

`} while ( !store_if_reserved( res, val ) );`

`*act = val;`

`return true;`
`}`

`/**`
`* Replace the value of res with val, only if res originally contains zero.`
`* @param res [IN/OUT] The data location`
`* @param val [IN] The value to be conditionally written.`
`* @return true iff *res contained 0, and val was written; otherwise false.`
`*/`
`static inline bool replace_if_zero( reserved_t* res, const value_t val ) {`

`value_t act;`
`return replace_if_match( res, (value_t)0, val, &act );`
`}`

### Layer 3: The Spinlock

Some spinlock implementations use a void return type. I use a boolean because it provides a common interface with other lock types, including a bounded version of the spinlock, which effectively times out after a specified number of attempts.

`typedef reserved_t spinlock_t;`

`enum {`
`UNLOCKED = 0x0,`
`LOCKED = 0x9999 /* Can be any non-zero value, 1 is also popular */`
`};`

`/**`
`* Take the spinlock.`
`* @note To avoid deadlock, treat spinlock like a critical section,`
`* holding the lock for only a few instructions at most.`
`* @param lock [IN/OUT] the location of the lock to be acquired.`
`* @return true always.`
`*/`
`bool spinlock_acquire( volatile spinlock_t* lock ) {`

`do{ /* spin */}while(!replace_if_zero( (reserved_t*)lock, LOCKED ) );`
`return true;`
`}`

`/**`
`* Take the spinlock, but only spin for the count provided.`
`* @note To avoid deadlock, treat spinlock like a critical section,`
`* holding the lock for only a few instructions at most. If that is done,`
`* this version should never be needed.`
`* @param lock [IN/OUT] the location of the lock to be acquired.`
`* @param attempts [IN] the maximum number of attempts to make.`
`* @return true iff the lock was acquired; otherwise false.`
`*/`
`bool spinlock_acquire_bounded( volatile spinlock_t* lock, unsigned int attempts ) {`
`bool result = false;`

`while ( attempts-- && !result ) {`
`result = replace_if_zero( (reserved_t*)lock, LOCKED );`
`}`

`return result;`
`}`

`/**`
`* Unlock the spinlock. Since the lock is already held, nobody else should be`
`* trying to manipulate it, therefore no synchronization is needed for the release.`
`*/`
`static inline bool spinlock_release( volatile spinlock_t* lock ) {`
`*lock = UNLOCKED;`
`return true;`
`}`

A spinlock assumes that whoever holds it is making progress and will release it within a few instructions. On a single-core processor, that assumption can fail without anyone doing anything wrong. Suppose a low-priority task holds the lock and a higher-priority task preempts it, then spins on the same lock. The holder cannot run while the spinner occupies the processor, so the spinner waits forever. On a multicore system the holder would be running on another core and would finish, which is why spinlocks are more at home there.

On a single core system, peer tasks are the safe case. When neither task preempts the other, either can be interrupted and resumed, and the ordinary unbounded spinlock works. The dangerous case is a lock shared between an ISR and a task. If the ISR fires while the task holds the lock, the task cannot resume until the ISR returns, and an ISR that spins waiting for it never returns. Therefore, an ISR on a single core should not attempt to lock, but should simply check the lock status, and make an informed decision about how to proceed.

A spinlock suits a critical section of a few instructions. Anything longer, or anything that might wait on something else, belongs under a lock that puts the waiter to sleep. Mutexes and semaphores do that through the scheduler. A waiting task yields the processor and is woken when the resource is released. The cost is a scheduler round trip at each end, which is a good trade for long holds and a poor one for short ones.

### Layer 4: The Data Structure Lock

The driver I was working with had a large data structure that was maintained by multiple peer tasks. In order to prevent corruption, each task had to obtain a lock on the structure before updating it.

A spinlock is the wrong tool for protecting a data structure while it is being modified, because that takes too long to spin through. A lock is meant for coarser-grained access. I used the spinlock for one narrow job, which was to make the acquisition of a lock safe. In this driver the spinlock was used to acquire two different locks.

The lock word is an ordinary variable, and reading it and then writing it has the same race that the first installment described. The spinlock closes that gap. Inside it, the function checks again that the lock is free, sets it, and releases the spinlock immediately. The spinlock is held for only those few instructions, while the lock itself can stay held for as long as the work requires.

`lock_acquire` never waits. When the lock is held, it returns false and leaves the decision to the caller. The initial check before the spinlock is optional, and as the comment says, it can be cheaper than acquiring the spinlock at all. Release is a plain store, for the same reason as before.

`typedef unsigned int lock_t;`

`/**`
`* Query if the lock is currently held.`
`* @param lock [IN] The lock to be checked.`
`* @return true iff lock is currently held; otherwise false.`
`*/`
`static inline bool lock_is_held( volatile lock_t* lock ) {`
`return (*lock != UNLOCKED);`
`}`

`/**`
`* Acquire the lock. Take the lock if it is available, otherwise fail.`
`* @param lock [IN/OUT] The lock to be held.`
`* @return true iff the lock was acquired; otherwise false.`
`*/`
`bool lock_acquire( volatile lock_t* lock ) {`

`static spinlock_t spinlock = UNLOCKED;`
`bool lockAcquired = false;`

`/* Pre-check is optional, but can be cheaper than the spinlock acquisition */`
`if ( *lock == UNLOCKED ) {`

`if ( spinlock_acquire( &spinlock ) ) {`
`/* Within the spinlock we can take the lock if it is available */`

`/* In case value changed before we entered the spinlock */`
`if (*lock == UNLOCKED) {`
`*lock = LOCKED;`
`lockAcquired = true; /* update return value */`
`}`

`spinlock_release( &spinlock );`
`}`
`}`

`return lockAcquired;`
`}`

`/**`
`* Unlock the lock. Since the lock is already held, nobody else should be`
`* trying to manipulate it, therefore no synchronization is needed for the release.`
`*/`
`bool lock_release( volatile lock_t* lock ) {`

`*lock = UNLOCKED;`

`return true;`
`}`

The above four layers provided the full range of locking needed for my driver port.

### Bonus Layer: The Keyed Data Structure Lock

The last layer adds a key. An ordinary lock can be released by anyone who can write to it, including a process that never acquired it. A keyed lock hands the acquirer a key, and only that key releases it.

`typedef unsigned int keylock_t;`
`typedef keylock_t keyval_t;`

`/*`
`* Similar to ordinary lock, but in order to prevent unlocking by unauthorized process,`
`* a key is issued by the acquire, and the same key is needed to perform the unlock operation.`
`*/`

`/**`
`* Query if the lock is currently held.`
`* @param lock [IN] The lock to be checked.`
`* @return true iff lock is currently held; otherwise false.`
`*/`
`static inline bool keylock_is_held( volatile keylock_t* lock ) {`
`return (*lock != UNLOCKED);`
`}`

`/**`
`* Acquire the lock using the provided key.`
`* @note key values equivalent to LOCKED and UNLOCKED are illegal,`
`* and will result in failed acquisition.`
`* @param lock [IN/OUT] The lock to be held.`
`* @param key [IN] The key value to be applied.`
`* @return The key value required to unlock iff the lock was acquired; otherwise 0.`
`*/``keyval_t keylock_acquire_set( volatile keylock_t* lock, const keyval_t key ) {`

`static spinlock_t spinlock = UNLOCKED;`
`keyval_t lockAcquired = (keyval_t)0;`

`/* Pre-check is cheaper than the spinlock acquisition */`
`if ( *lock == UNLOCKED && key != UNLOCKED && key != LOCKED ) {`
`if ( spinlock_acquire( &spinlock ) ) {`
`/* Within the spinlock we can take the lock if it is available */`
`/* In case value changed before we entered the spinlock */`
`if (*lock == UNLOCKED) {`
`*lock = (value_t)(LOCKED ^ (value_t)key);`
`lockAcquired = key;`
`}`
`spinlock_release( &spinlock );`
`}`
`}`
`return lockAcquired;`
`}`

`/**`
`* Get a fresh key for a keylock. This is a super simple implementation,`
`* but an even simpler (non-portable) one would capture the timestamp and`
`* convert it to a key.`
`* @return the value of the new key.`
`*/`
`keyval_t keylock_get_key( void ) {`
`static keyval_t key = 1;`
`static spinlock_t lock = UNLOCKED;`
`keyval_t result = 0;`

`do {`
`if( spinlock_acquire( &lock ) )``{`
`result = ++key;`
`spinlock_release( &lock );`
`}`
`} while( result == LOCKED || result == UNLOCKED );`
`return result;`
`}`

`/**`
`* Acquire the lock using the provided key.`
`* @note key values equivalent to LOCKED and UNLOCKED are illegal,`
`* and will result in failed acquisition.`
`* @param lock [IN/OUT] The lock to be held.`
`* @return The key value required to unlock iff the lock was acquired; otherwise 0.`
`*/`

`keyval_t keylock_acquire_autokey( volatile keylock_t* lock ) {`
`return keylock_acquire_set( lock, keylock_get_key() );`
`}`

`/**`
`* Unlock the lock if the correct key is supplied. Since the lock is already held,`
`* nobody else should be trying to manipulate it, therefore no synchronization is`
`* needed for the release.`

`* @param lock [IN/OUT] The lock to be released.`
`* @param key [IN] The key value to be applied.`
`*/`

`bool keylock_release( volatile keylock_t* lock, const keyval_t key ) {`

`if ( LOCKED == (*lock ^ key ) ) {`
`*lock = UNLOCKED;`
`return true;`
`}`
`return false;`
`}`

The lock word holds the locked value combined with the key by exclusive-or, and release recomputes the combination from the key supplied. The acquire functions return the key on success and zero on failure. The automatic version generates keys itself, so the caller never has to choose one.

I never needed the keylock in the driver, but I did consider it, and the idea is useful in safety-critical work. Picture processes of different criticality sharing a common datastore behind a protected API. The safety assumption is that a lower-criticality process could go rogue and make an arbitrary call that unlocks the datastore while a higher-criticality process is using it, and the result is corruption. With a keyed lock, a process that doesn’t hold the key to a lock cannot release it. Keys could be fixed, assigned per-process, or they could be issued automatically at lock acquisition, and the above code accommodates either method. By itself, it does not provide total safety, but it adds a layer of protection.

Most teams with a gap like this take one of two paths. They avoid assembly entirely and restructure their code around whatever the RTOS provides, or they build a large assembly-heavy subsystem. The approach in this series sits between the two. A small, deliberate assembly foundation supports everything else, and the logic that is hard to get right is written in C, in the language the team reads.

What still strikes me is how little code is involved. The whole stack, from derived atomics to keyed locks, comes to a few dozen lines of C on top of a few wrapped assembly language opcodes. None of the algorithms is complicated, and each can be understood in one sitting. Together they solve problems that a pure C approach cannot, without undue confusion that often accompanies assembly language.

If you'd like to take a closer look at the code and possibly try it out; feel free to download the [sample code](https://safecodetech.com/insights/articles/downloadables/download/4-c/1-locks-in-layers.html)that was created for these articles. It's been tested to compile on ARM; but this is fresh code created from the original concepts; so the runtime testing is all on you ;).
