In a multithreaded program that uses condition variables, after one thread calls wait, when no other thread has called notify,
the original waiting thread just restarts by itself. Truly unreasonable.
How the Problem Arose
I first encountered Spurious wakeup in Chen Shuo’s “Linux Multithreaded Server Programming”. Later, while helping my sister modify a server,
I also ran into condition-variable usage problems. Without question, half-understanding cannot produce excellent programs.
Below is what my program wanted to implement: a mechanism similar to garbage collection.
- The main thread needs to know the variables to be cleaned up (stored in a
vector<int>); afterwards the main thread traverses thevectorto do the cleanup.
- The collector thread always traverses all or part of the objects, stores the variables to clean into the
vector, then suspends until the main thread finishes cleaning before proceeding to the next step.
Why not put the object traversal in the main thread too?
A: The main reason is that traversing all objects takes too long; the main thread cannot accept that time requirement.
Why not put the collection activity in the child thread either?
A: When the objects needing cleanup are always numerous and block the main thread’s execution, consider moving it into the child thread. But designing it that way further increases the program’s complexity. If there’s such a need, I’ll open another post to analyze solutions.
First Version of the Program
- Shared variables
1 | class Server { |
- Main thread
1 | while(1) { |
- Collector thread
1 | while(1) { |
Research and Analysis
The man pages
man 3 pthread_cond_wait
Spurious wakeups from the
pthread_cond_timedwait()orpthread_cond_wait()functions may occur. Since the return frompthread_cond_timedwait()orpthread_cond_wait()does not imply anything about the value of this predicate, the predicate should be re-evaluated upon such return.
We called the cond_wait function; it always returns, but returning does not indicate that the predicate (bool value) has changed, so
the real state at return must be evaluated again and again.
Why, knowing this problem exists, is it only documented instead of being fixed in the program?
From the (POSIX Thread Architect) email:
- Religiously using a loop to check the predicate is indeed good coding practice
- It’s not hard to imagine that if user code adds the checking, the underlying layer can optimize the synchronization mechanism to improve performance
Do Spurious wakeups Really Happen?
Many blogs in the answers are no longer reachable; one friend said 20% of wakeups are Spurious wakeups. In the future,
if I get the chance, I will modify lots of production code so it can detect Spurious wakeups.
Test Program
The program above made a very serious mistake — a textbook error. Thanks to my sister for pointing it out earnestly:
a condition variable must be used together with a mutex — that is, perform wait/signal operations only after locking.
Final Implementation
- Shared variables
1 | class Server { |
- Main thread
1 | while(1) { |
- Collector thread
1 | while(1) { |
Unfortunately, I ran the program many times and still never saw the
Spurious wakeupphenomenon appear. After I start working, I will try hard to observe it in practice. Below is the test code.
Future Notes
Remember three sentences religiously — textbook-style (for the wait side):
- It must be used together with a mutex; reading and writing the boolean expression must be protected by the mutex
- wait may only be called while the mutex is locked
- Put the boolean condition check and wait inside a while loop