SQLite carried a corruption bug for 16 years. Here's why yours is probably fine, and what to check anyway.
This week Tailscale published a postmortem that every team running SQLite should read. For six months, starting last August, their control plane kept corrupting databases. Nineteen separate incidents, no obvious trigger, no recent code change that explained it. The cause turned out to be a data race that had been sitting in SQLite since version 3.7.0, which shipped in July 2010. It hid there for roughly sixteen years.
The SQLite team calls it the WAL-reset bug. It only bites databases running in WAL (write-ahead logging) mode when two or more connections have the same file open across separate threads or processes, and they try to write or checkpoint at the same instant. When the timing lines up, a checkpoint can leave a flag in the WAL index claiming that part of the log was already copied into the main database file when it wasn't. A later checkpoint then skips that data, and the file goes corrupt. It was fixed in SQLite 3.51.3 on March 13, 2026, with backports to 3.44.6 and 3.50.7.
Here's the part worth sitting with. The SQLite developers could not reproduce this bug on purpose. They had to patch their own source to fire a callback at precisely the wrong moment during a checkpoint. Their own telemetry puts the real-world hit rate at or below the rate of SSD failures and cosmic-ray bit flips. So if you run SQLite in a normal configuration, you were almost certainly never affected, and you still probably aren't. Tailscale got hit because they took manual control of checkpointing and ran it aggressively for fast backups. They stepped off the well-worn path, and a one-in-a-billion race became a monthly event.
What we'd actually do about it
Upgrade, but don't panic. Get to SQLite 3.51.3 or newer, or one of the backports, on your own schedule. The catch is that "your SQLite version" is rarely obvious. SQLite is embedded almost everywhere, and you don't install it directly. It ships inside your language runtime, your ORM, your mobile OS, and your edge platform.
In our native iOS and Android work, this is the easy one to forget. The system SQLite is tied to the OS version your users run, not the version you built against, and you don't get to pick when they upgrade. If your app does anything unusual with connections or checkpoints, say a background sync thread writing while the UI reads, bundle your own SQLite build so you control the version instead of inheriting whatever the device ships. We've done exactly that in offline-first healthcare and IoT apps, where a corrupt local store means data the user can never get back.
On the backend and edge side, check what your driver actually links against. From our PHP and Docker work, the SQLite version baked into a base image can lag months behind upstream, and "we're on the latest framework" tells you nothing about the C library underneath. Serverless and edge SQLite, meaning D1, Turso, LiteFS and friends, is worth a specific look, because those platforms lean hard on WAL mode and concurrent access. That's the exact shape this bug needs.
The bigger lesson is about backups
The detail that stuck with me isn't the race condition. It's that Tailscale only noticed because a separate pipeline read their S3 backups and ran PRAGMA integrity_check against them. Their live databases looked healthy. The corruption was already sitting in the files they'd been storing for later.
During our security and pentesting engagements, backups are where we find the scariest gaps. Teams take them religiously and never once restore them. A backup you have not restored is a hypothesis, not a safety net. If you store SQLite files, run PRAGMA integrity_check on a copy on a schedule, and actually restore into a scratch environment now and then. Corruption that only shows up the day you need the backup is the worst way to learn this.
One more thing from the writeup that's easy to miss. After Tailscale deployed the fix, a rounding change in the release triggered false corruption warnings on databases that used expression indexes. They caught it because they rolled out to a few canary shards first, then the rest. Phased rollouts aren't bureaucracy. They're how you find the second bug that the fix for the first bug introduced.
What you can do this week
Find out which SQLite version your production systems and your shipped apps are really running. Not the framework version, the actual SQLite library. Then pick one backup and prove you can restore it clean. That's an afternoon of work, and for most teams it's worth more than the upgrade itself.
If chasing down which SQLite version your apps and services actually run, and whether your backups would survive a restore, sounds familiar, let's talk.