Pi cannot be read as another linear JSONL stream. Its history is a tree with a durable leaf, orphan roots, branch summaries, and two compaction forms, so the active context is something the format states rather than something line order implies. The adapter keeps those semantics inside itself and projects the result into the existing canonical tables. Sessions are keyed by (normalized header cwd, header id) rather than by path, because Pi's --session-id lookup is project-local: two projects may reuse an id, while a move or an identical copy is still one session. Discovery covers both layouts Pi writes and fingerprints each file by mtime, ctime, size and inode, so a rewrite that preserves mtime is not read as unchanged. Abandoned branches are preserved rather than dropped. Visibility becomes three-state -- visible, inactive, hidden -- and helpers return only visible rows until includeInactive asks for the superseded path, labeling every row so a caller knows which it holds. Usage counts all three, because an abandoned call still spent tokens; message_count reports only the visible transcript. A committed MIT-licensed oracle transcribed from Pi 0.83.0 pins the context algorithms, and a fixed-seed differential runs 512 generated sessions against it on every test run. Schema changes are additive.
5.6 KiB
Indexing is a registry of pure provider adapters over one shared persist layer
Revised 2026-07-08. The first draft framed the parse layer as a single "parse core" with "two thin persist layers, one per binding." That was wrong on both axes and is corrected below: the parse layer is a registry of per-provider adapters (driven by the multi-provider roadmap), and there is one shared persist layer, not one per binding.
Context. Obelisk had two divergent full indexers — the former
scripts/indexer.mjs (node:sqlite, the former skill-embedded runtime) and
app/indexer.js
(better-sqlite3, Electron
app) — that duplicated the same Claude and Codex JSONL parsing and had silently
diverged in write semantics (INSERT OR REPLACE vs ON CONFLICT DO UPDATE,
message-count accumulation). Two forces shape the fix: (1) the roadmap will add
more transcript sources — opencode, pi, and others — so the parse layer must be
pluggable, not one monolith; (2) node:sqlite and better-sqlite3 share the
same prepare/run/get/all API, so persistence is already nearly
binding-agnostic and does not need a per-binding implementation.
Decision. Split indexing along two orthogonal axes.
- Provider axis — a registry of pure adapters. Each source (Claude Code,
Codex, Kimi Code, Pi, …) is a provider adapter implementing one complete
boundary: serializable descriptor metadata,
watchRoots(root),discover(context) → IndexUnit[],parse(unit, cursor) → Iterable<Record>, andraw(lookup). AnIndexUnitis deliberately not a file abstraction: Kimi uses one session directory containing state plus multiple agent wire logs. An adapter is pure: it emits normalized records and never touches a database. Discovery receives a read-only view of the provider's already-indexed session paths. A full-reparse adapter can therefore attachretractSessionIdsto a replacement or tombstone unit without querying SQLite itself. Retraction and replacement records commit in the same unit transaction, so a failed parse preserves the last complete snapshot. Provider identity may therefore be richer than a wire-level ID. Pi, whose explicit session IDs are project-local, deterministically namespaces the header ID by the normalized header cwd; source paths remain provenance rather than identity, so copying or moving a transcript does not rename the session. Adding a source means adding one adapter and registering it; nothing else changes.parseexposes an iterator as its common interface and streams when the provider semantics permit it. An adapter may buffer one completeIndexUnitwhen correctness requires whole-unit semantics — for example, Codex duplicate reconciliation, Kimicontext.undo/context.clearreplay, or Pi tree projection. Each adapter maps its own resume/change semantics onto an opaque cursor stored exactly inindex_state.cursor;mtimeandlines_processedremain compatibility/index-inspection columns. A provider-wide replay is a destructive snapshot boundary. Discovery reports any source location it could not enumerate; destructive rebuilds fail closed when such a report exists. Cleanup, parsing, persistence, FTS rebuild and marker publication commit in one transaction, so a parse failure cannot replace the last-good index. The emittedTranscriptRecordstream is also the input to provider-independent session detail assembly; see ADR-0007. - Persist axis — one shared orchestration. A single provider-agnostic,
binding-agnostic layer consumes records from any adapter and writes them:
incremental
index_statebookkeeping, FTS maintenance, and the canonical upsert (ON CONFLICT(uuid) DO UPDATE) write semantics reconciled from the drift on 2026-07-08. The database handle is injected, sonode:sqlite(CLI) andbetter-sqlite3(app) run the same code — there is no per-binding persist layer.
Two indexing modes share all of the above and differ only in trigger:
daemon mode (the app, and potentially a future CLI daemon, watches and keeps
the index fresh) and passive pull mode (a CLI command indexes on invocation
when no daemon is active). They never write concurrently — passive mode detects
a fresh daemon via heartbeat markers in index_state (daemon arbitration).
Consequences. Golden tests anchor on each adapter's parse output (feed
fixture JSONL, assert the yielded record sequence) — independent of binding and
persistence. The app's richer changed-path discovery becomes a discover
strategy injected into the shared orchestration, not a fork of it. The Electron
main process migrates to ESM (ADR-0003) to import the shared core. The real work
is disentangling the currently interleaved parse-and-write inside indexJsonl /
indexCodexJsonl into (pure adapter parse) + (shared persist).
The normalized TranscriptRecord union is the stable center of the design, and
SQLite is one serialization adapter for it. Provider-only concepts are either
projected lossily into that language or ignored. A genuinely shared concept may
extend the canonical language by explicit decision: summary-generating model
calls, for example, carry the same normalized input/output usage as ordinary
messages. The registry, not provider switches, drives both indexers, watcher
roots, persisted source roots, source catalog/UI labels and colors, and
raw-record routing. Adding Pi therefore adds no Pi-specific branch to the
shared schema, persist layer, indexers, settings, query API, or renderer; the
provenance and summary-usage additions are provider-neutral contracts.