Architecture

The evidence-locked extraction architecture and its planned product phases.

Migrated from docs/roadmap-transcript-intelligence.md

Roadmap: Transcript Intelligence

Status: draft, 2026-07-23. Owner decision pending on phase order.

Vision

One nightly, evidence-locked pipeline that reads the user's real session transcripts and turns them into truth-anchored products: goal reconciliation, highlights digests, decision journal, frustration analytics, company-context proposals, product/feature discovery, and hook proposals.

Not ten separate systems — one extraction engine, one append-only event log, many views.

Foundations already in place (as of 2026-07-23)

  • scripts/report-goals-completed-today.mjs — the transcript-truth oracle: mtime-filtered session scan, chunked LLM candidate extraction, decision pass with conservative coercion, evidence quotes + provenance, model fallback chain (kimi/k3codex/gpt-5.5kimi/kimi-for-coding), content-keyed caches, --apply mode closing proven completions.
  • oko-cli auto-goals reconcile-today — once-per-day scheduling at the predicted end-of-day (WorkTelemetry.forecast); detached from suggestions tick.
  • Tests/OkoKitTests/SlackGoalScenarioPolicyTests.swift — the 3 live truth-verifying Slack scenario tests (zero seeding).
  • Existing OkoKit pieces: FrustrationDetector, StatsAggregator, TranscriptIndexer (SQL), CompanyContextExtractor, OkoInteractionRequestStore (approval flow), work_costs.json.

Architecture (target)

transcripts (~/.omp, ~/.claude, ~/.kimi-code)
      │
      ▼
extractors/            ← one typed extractor per signal class
  goals / highlights / decisions / frustrations /
  patterns / context-facts / products / hooks
      │  evidence-locked: quote + provenance; unverifiable → unknown
      ▼
event log (append-only, ~/.oko/intelligence/events.jsonl or sqlite)
      │
      ▼
views                ← thin readers of the log
  Slack digest / app panes / approval requests / snapshot writes

Rules inherited from the oracle: truth comes from user-authored evidence only; assistant claims never prove completion; every surfaced claim carries its evidence quote; anything mutating state beyond goal-closure goes through the approval flow.

Phase 0 — done (2026-07-23)

  • Zero-seeding live tests (the 3 Slack scenarios) verifying truth end-to-end.
  • Oracle hardened for real-world data; --apply goal reconciliation cron.

Phase 1 — Highlights digest + decision journal

Smallest step with immediate daily value; reuses the oracle engine directly.

  • Extractor: highlights — user-acknowledged completions, shipped work, deployments, merges; plus decisions — verbatim user decisions with timestamps ("zero zasiewania", "tylko 3 testy", …).
  • Event log v0: ~/.oko/intelligence/events-<day>.jsonl.
  • View: end-of-day Slack digest (after reconcile-today, same scheduler): "Co się dziś naprawdę stało" — highlights + decisions, each with evidence.
  • Verification: extend test 1's philosophy — a scenario test that seeds nothing and diffs the digest against oracle-provable events.

Phase 2 — Frustration analytics

  • Extend FrustrationDetector signals with transcript-level patterns: repeated re-asks, scope corrections, undo commands, "nie rozumiem" class.
  • Cluster root causes (agent ignored instruction / tool failure / latency / wrong scope) — weekly friction report per feature area.
  • Feeds Phase 5 (hook proposals) with the top friction sources.

Phase 3 — Company context proposals

  • Extractor: new facts (repos, services, tables, people, deploy targets).
  • View: approval requests via OkoInteractionRequestStore — user approves edits to the company context store; nothing writes blindly.

Phase 4 — Product & feature discovery

  • Project genesis detection: new repo/table/plist/app appearing in sessions → proposal to register in the product portfolio.
  • Recurring manual workflows → automation candidates.
  • "Marzy mi się…" class: feature wishes uttered in passing → backlog proposals through the approval flow.

Phase 5 — Hook & automation proposals

  • Recurring user corrections of agent behavior → hook proposals (the manual equivalent already exists as auto-added rules in CLAUDE.md).
  • Dangerous-operation near-misses → guard-hook proposals.
  • Delivered as approval requests with the evidence trail attached.

Phase 6 — Cost-vs-outcome and model routing

  • Join work_costs.json with truth verdicts: cost per actually-completed goal, per model/agent.
  • Feed the model-router with per-task-class success rates.

Cross-cutting

  • Every phase: extractor emits evidence-locked events; no evidence → no event.
  • Scheduling: all extractors run inside the same nightly window as reconcile-today (one scan, many extractors — the 5164-file scan happens once, extractors share snippets).
  • Caching: content-keyed per extractor, like the oracle's chunk caches.
  • Privacy: local-user scope only; org-wide aggregates never leave the machine without the approval flow.

Phase 7 — Viral features (transcript-native, receipt-backed)

The differentiator against Spotify-Wrapped clones: every stat, achievement and roast item carries its receipt — a transcript quote + session/line provenance, verified by the same oracle contract. "Pics or it didn't happen" is built in.

Oko Wrapped (rok w review)

Shareable card pack (HTML→PNG) from the event log: total prompts, tokens, cost, top model, longest session, latest night, most-said phrase, goals opened vs closed ratio, busiest day heatmap (GitHub-style contribution graph of prompts/completions).

Percentile flex ("top 1%")

Anonymous percentile cards: tokens/day, parallel agents, completed goals, prompt-to-ship ratio. Scope: org pool or global anonymous pool (opt-in only). Card: "Top 0.4% — 41 agents w ciągu jednego dnia. Dowód: 41 sesji, pierwsza 06:12, ostatnia 03:58."

Achievements (hardcore)

  • Syzyf — ten sam test/build padł 10+ razy w jednej sesji, a i tak wysłałeś. Receipt: sekwencja fail→fix z timestampem.
  • Prompt snajper — cel domknięty jednym promptem, zero follow-upów.
  • Klub 3:00 — produktywna sesja po 3 w nocy (streak wariant: tydzień).
  • YOLO — deploy/push w piątek po 22:00.
  • Cofacz — rekord "cofnij" na dzień (counter + najlepszy cytat).
  • Pogromca bookkeeping-u — udowodniłeś, że snapshot kłamał (oracle discrepancy > 10 celów).
  • Token-żarłok — 1M+ tokenów w 24h z kosztem w $ jako subtytuł.
  • Płomiennica/Piekarz — streak N dni z rzędu z domkniętym celem.
  • Full-stack człowiek — 5+ różnych agentów/modeli w jednym dniu.
  • Wieża Babel — rozkazywałeś agentom w 3+ językach w jednej sesji.

Roast mode (opt-in)

LLM roast card, ale każdy punkt ma dowód: najczęściej powtarzana komenda ("zrob commit and push" ×47), cel otwierany co tydzień od miesiąca, porzucone sesje z ostatnim zdaniem bez odpowiedzi, "to już ostatnia rzecz" ×N, najdłuższa tirada frustracji z cytatem. Format: 6-10 bullets + werdykt ("Promptuje jak senior, commituje jak stażysta"). Never org-visible.

Rage bingo

Karta 5×5 frustracji z dnia/tygodnia: "powiedziałem 'nie rozumiem' 3×", "agent zrobił coś o co nie prosiłem", "cofnij wszystko". Bingo = roast card.

Boss battle (RPG)

Cele jako questy: XP za domknięcie (wagą = koszt tokenów), poziom postaci, boss = najdłużej otwarty cel ("Write sign-in-controls e2e spec — 23 dni, 3421 HP"). Weekly raid report na Slacka.

Turing tax

Live counter: ile $ poszło dziś w tokeny vs ile celów faktycznie domknięto (prawda z oracle). Dzienny werdykt: "Zapłaciłeś $31.20 za 2 cele. Jeden to był 'zrob commit'."

Ship/confidence market

Rano: predykcja ile z dzisiejszych celów zamkniesz. Wieczorem: rozliczenie z prawdą. Streak trafień jako achievement.

Zasady

  • Opt-in per feature, local-first; karty eksportowane dopiero na życzenie.
  • Org/team comparisons wyłącznie przez approval flow, zanonimizowane.
  • Każdy element ma evidence link — to jest cały żart i cały produkt.