Files
obelisk/docs/adr/0007-canonical-transcript-session-detail-seam.md
T
SaladDay 2589384e68 feat(core): add first-class Pi session indexing (#23)
Pi cannot be read as another linear JSONL stream. Its history is a tree with
a durable leaf, orphan roots, branch summaries, and two compaction forms, so
the active context is something the format states rather than something line
order implies. The adapter keeps those semantics inside itself and projects
the result into the existing canonical tables.

Sessions are keyed by (normalized header cwd, header id) rather than by path,
because Pi's --session-id lookup is project-local: two projects may reuse an
id, while a move or an identical copy is still one session. Discovery covers
both layouts Pi writes and fingerprints each file by mtime, ctime, size and
inode, so a rewrite that preserves mtime is not read as unchanged.

Abandoned branches are preserved rather than dropped. Visibility becomes
three-state -- visible, inactive, hidden -- and helpers return only visible
rows until includeInactive asks for the superseded path, labeling every row
so a caller knows which it holds. Usage counts all three, because an
abandoned call still spent tokens; message_count reports only the visible
transcript.

A committed MIT-licensed oracle transcribed from Pi 0.83.0 pins the context
algorithms, and a fixed-seed differential runs 512 generated sessions against
it on every test run. Schema changes are additive.
2026-08-04 23:33:01 +08:00

3.4 KiB

Canonical transcript records are the session-detail seam

Context. Provider adapters originally emitted database-shaped records, while the desktop app reconstructed presentation semantics after querying SQLite. Although that reconstruction had no explicit provider switch, it still inferred metadata from raw message text. As more providers are added, those heuristics would make provider semantics leak into a shared presentation module and allow the direct parse path to drift from the persisted path.

Decision. Every provider adapter emits a canonical TranscriptRecord stream. The adapter owns all source-specific interpretation: duplicate raw events, stable identities, tool relationships, message classification, and visibility. Messages and summaries carry provider-normalized visibility. visibility is separate from is_meta. visible is current evidence. inactive is physical evidence that the provider explicitly attests was superseded; default queries omit it, and supported helpers may return it only with includeInactive: true. hidden is display-suppressed or transport-only material and no standard helper returns it. Session detail and the desktop app remain visible-only, while visible system evidence can remain a metadata card. Presentation-sensitive concepts are explicit canonical fields: tool calls carry a presentation class, Skill instructions carry a content type, and workflows carry their parent tool-call identity. Summaries carry normalized input/output usage when their provider performed a separate model call; cached input is folded into input usage by the provider, as it is for messages. Visibility does not erase accounting: aggregate usage includes model calls that were later abandoned.

Provider capability determines whether inactive is meaningful. Pi attests supersession through branch, leaf, and compaction state. Kimi undo/clear can attest it, but preserving that history is separate work. Claude transcripts do not attest rewind or current-leaf state, and Codex sessions do not branch.

The Core assembleSessionDetail(input) module is the only session-detail seam. It accepts either a provider's complete transcript stream from a fresh parse (cursor = null) or table-shaped rows after a persistence round-trip. A delta parse cannot produce a complete detail snapshot without prior state, so the assembler rejects a SessionRecord whose countMode is delta; incremental UI updates use the existing snapshot/patch seam. Its internal row adapter restores the canonical record language before assembly. The implementation may sort, group thinking, and attach tool results, subagents, and workflows, but it never checks the provider and never parses message text to recover provider semantics. Tool names are likewise display data, not assembly control flow.

The persist layer only serializes transcript records and cursor state. SQLite is not the source of transcript semantics, and a persistence round-trip must not change the assembled result.

Consequences. A new provider is complete only when its canonical transcript can pass directly through assembleSessionDetail. Provider conformance tests cover that seam, while persistence tests verify that canonical classification survives a database round-trip. Codex-owned normalization now classifies hidden context envelopes and structurally removes image wrappers before duplicate reconciliation. Adding another provider does not add branches to the app's session-detail code.