First-Person Developer POV: Server Infrastructure & Sync Hardware
Let’s talk about the absolute worst bug in client-side software: the silent data overwrite.
Imagine this scenario:
- You open your notes app on your iPhone while waiting in line for a morning pour-over. You have no cellular signal. You add a critical item to your daily checklist.
- Simultaneously, your laptop at home is open on your desk, and your scheduled background backup triggers a timestamp update on the exact same note.
- When your phone reconnects to Wi-Fi ten minutes later, a naive "Last-Write-Wins" (LWW) sync algorithm compares the two timestamps:
- Phone timestamp:
10:14:02.105 - Laptop timestamp:
10:14:03.450 - The server looks at the timestamps, shrugs, declares the laptop write the "winner," and silently deletes the note you just wrote on your phone.
No error prompt. No merge dialogue. Just quiet, irreversible data loss.
For years, web developers treated multi-master offline synchronization as a terrifying, unsolvable dark art.
Enter Conflict-free Replicated Data Types (CRDTs)—the mathematical framework that makes distributed state reconciliation provably deterministic.
The Two Churches of CRDTs
Mathematically, CRDTs fall into two distinct families:
1. State-Based CRDTs (CvRDT - Convergent)
In a state-based architecture, whenever a local mutation occurs, the node transmits its entire state (or a compressed delta of changes) to its peers. The reconciliation function is a monotonic join-semilattice:
$$S_{\text{new}} = S_A \sqcup S_B$$
To work correctly, the join function must satisfy three mathematical properties:
- Commutative: Merging A into B produces the exact same result as merging B into A ($A \sqcup B = B \sqcup A$).
- Associative: Grouping doesn't matter ($(A \sqcup B) \sqcup C = A \sqcup (B \sqcup C)$).
- Idempotent: Merging the same state twice changes nothing ($A \sqcup A = A$).
2. Operation-Based CRDTs (CmRDT - Commutative)
Instead of sending state payloads, nodes transmit atomic mutation operations across an exactly-once causal broadcast channel. As long as concurrent operations commute, all nodes deterministically arrive at the identical state.
For relational client applications (notes, tasks, tags), delta-based state CRDTs combined with Hybrid Logical Clocks provide the sweet spot of reliability over unreliable consumer mobile networks.
Wall-Clock Time Is a Lie: The Need for HLCs
The biggest trap in distributed systems is trusting physical system clocks.
Your iPhone’s clock and your MacBook’s clock are not synchronized. Network Time Protocol (NTP) adjustments routinely cause physical system clocks to skew by hundreds of milliseconds, or worse, jump backward in time during a sync correction.
If you rely on Date.now() to order distributed events, causality breaks. You can literally receive an "edit" that appears to have occurred five minutes before the document was created.
A Hybrid Logical Clock (HLC) combines the physical system time with a monotonic logical counter:
interface HLC {
physicalTime: number; // Unix timestamp in milliseconds
logicalCounter: number; // Increments if events happen within the same ms
nodeId: string; // Unique peer identifier (UUIDv4)
}
export function advanceHLC(current: HLC, receivedPhysicalTime: number): HLC {
const now = Date.now();
const maxPhysical = Math.max(current.physicalTime, receivedPhysicalTime, now);
// If time hasn't advanced physically, increment the logical counter
if (maxPhysical === current.physicalTime) {
return {
physicalTime: current.physicalTime,
logicalCounter: current.logicalCounter + 1,
nodeId: current.nodeId
};
}
// Otherwise, advance physical time and reset the counter
return {
physicalTime: maxPhysical,
logicalCounter: 0,
nodeId: current.nodeId
};
}With an HLC, causality is strictly preserved. Every single mutation across all devices receives a deterministic, strictly monotonic ordering that never jumps backward.
The Tombstone Problem (and How to Clean It Up)
In a traditional centralized database, deleting a row is trivial: DELETE FROM notes WHERE id = ?.
In a distributed local-first system, if Device A deletes a note by simply deleting the local SQLite row, Device B has no way of knowing it was deleted. The next time Device B syncs, it will see the note on its local disk, conclude that Device A is missing this record, and resurrect the deleted note from the dead.
To prevent zombie resurrection, deletions must be tracked as tombstones:
ALTER TABLE notes ADD COLUMN is_deleted INTEGER DEFAULT 0;
ALTER TABLE notes ADD COLUMN deleted_at TEXT;When a user deletes a note, you set is_deleted = 1 and broadcast the tombstone mutation. Every peer marks the record as deleted in their local SQLite table.
The 30-Day Compaction Horizon
Of course, if you never delete tombstones, your local SQLite database will grow indefinitely with ghost records.
The solution is establishing a Compaction Horizon (typically 30 days). Once all registered devices have acknowledged receiving changes older than 30 days, a background SQLite vacuum job safely purges the tombstone rows forever:
DELETE FROM notes
WHERE is_deleted = 1
AND deleted_at < datetime('now', '-30 days');Elena's Unfiltered Take
Building local-first sync with CRDTs requires unlearning twenty years of centralized database habits. You have to stop thinking in terms of locking rows and start thinking in terms of mathematical convergence.
Once you cross that conceptual bridge, you can never go back. Your applications become resilient, permanent, and impervious to network hiccups. That's real engineering.