Skip to main content

PortfolioCase study

RC-01Live

RailCite

Research on track.

A trust-first assistant for Indian Railways commercial circulars: it cites the right circular with number, date and supersession lineage — or refuses.

Illustration of a cream streamliner with a rust chevron coming down the line at dusk, the sun setting behind it — every third sleeper a ruled document page, a signal ahead showing green, a lit signal box, and a stack of bound volumes with a magnifier in the foreground.
  • 5,760

    documents indexed

    Measuredas of 15 Sep 2026

    live corpus; a nightly crawl keeps adding circulars

  • 0

    invented citations

    Structural

    by construction: the validator drops any citation that doesn't resolve

  • 68%

    of ingested PDFs needed OCR

    Measuredas of 7 Sep 2026

    3,865 of 5,687 ingested PDFs

Evidence keyMeasuredStructural

The problem

One wrong circular can damage an inspector’s credibility.

A Chief Commercial Inspector has to defend every demurrage or wharfage decision — across years of circulars, scanned PDFs and rules that were quietly superseded.

How a justification gets written today
  1. Yearly PDF lists
  2. Scanned circulars
  3. Guess the current version
  4. Check supersession
  5. Write the justification
  6. Cite by hand

Instructions not included “should not be deemed to have been superseded simply because of their non-inclusion.”

— Railway Board Master Circular caveat, quoted in the Discovery PRD

The product

Ask, get the right rule — or a clear refusal.

The inspector asks in plain language. RailCite answers only from passages that govern the case, and says so when none does.

Illustration of a desk of ruled circulars under a magnifying glass, one page ticked as cited.
RailCite's Ask console: a demurrage-waiver case in the Commercial domain, answered “Cited from 8 passages in the Goods manual” with an inline citation marker [1] and an English/Hindi toggle.
A case answered from the Goods manual, cited inline
RailCite's Sources panel: numbered primary-circular cards with circular numbers, dates and page ranges, each marked “Verified text”.
Every citation opens its circular, page by page
  • Answered

    • The governing circular, with number and date
    • Supersession lineage: which version governs today
    • Every citation resolves to a real source
  • Refused

    • No passage governs the case
    • Designed as a success state, never an error

Product decisions

Refuse rather than fabricate.

  • Could have
    Always answer with the best-matching passage
    Chose
    Cite-or-refuse: refusal is a first-class success
    Because
    A confident wrong citation damages the inspector’s credibility, not the tool’s.
  • Could have
    Trust the prompt to behave
    Chose
    Validate every citation in code
    Because
    A validator drops any answer block whose citations don’t resolve; if every block drops, the answer becomes a refusal.
  • Could have
    Make cross-domain bleed unlikely
    Chose
    Make it impossible with a hard domain filter
    Because
    One commodity’s circular must never reach another’s case: “bleed has to be impossible, not merely unlikely.”

How trust works

Trust lives in the system, not just the prompt.

  1. Questiona signed-in request
  2. Embed + classifyVoyage-3 vectors; Claude Haiku routes the domain
  3. Retrievetop 8 passages, hard domain filter
  4. Thresholdcalibrated 0.45 → 0.32
  5. SynthesisClaude Sonnet 5, extractive, answered | refused
  6. Citation validatordrops citations that don’t resolve
  7. Answer or refusalwith supersession lineage
RailCite’s query pipeline, in seven steps

Evidence

What the records show — and what they don’t.

The retrieval threshold was calibrated against real queries rather than guessed, and the corpus runs live. Demand-side proof doesn’t exist yet.

  • 0.32

    calibrated relevance threshold

    Measured

    5 relevant + 3 irrelevant queries: irrelevant ≤ 0.25, relevant 0.29–0.66; a nonsense query was refused

  • 14,406

    chunks indexed, live

    Measuredas of 15 Sep 2026

  • 193

    supersession lineage links

    Measuredas of 7 Sep 2026

  • 22/40

    UX critique of the Ask screen

    Measuredas of 31 Aug 2026

    one P0, “Flagship starter refuses”; no fix is recorded

What isn’t measured

  • No usage data and no measured time-to-cited-answer.
  • No groundedness, retrieval-precision or latency eval; citation validity is structural, not sampled.
  • Sessions with real inspectors happened but weren’t logged.

Key learnings

Three lessons that made RailCite trustworthy.

  • Refusal is a feature

    A confident wrong citation is worse than no answer.

  • Trust belongs in code

    Citation validation and supersession checks enforce correctness beyond prompting.

  • Freshness is correctness

    A superseding circular turns a perfectly cited answer wrong — so the corpus is crawled nightly.

Sources

Everything here is backed by a real artifact.

  • PRD
  • Design
  • Architecture
  • Code
  • Evaluation
  • Test run
  • Live data
  • Build ledger

Evidence behind the RailCite case study

Private working documents are listed with what they support; only public sources link out.

  1. Discovery PRD

    PRD28 Aug 2026

    Supports: The problem, the “Ravi” persona and the Master Circular insight

  2. Final PRD

    PRD7 Sep 2026

    Supports: Ingest figures (5,687 ingested, 68% OCR, 193 lineage links) and the bleed fix

  3. Design North Star

    Design

    Supports: “Refuse is a first-class success state, never an error”

  4. Query pipeline

    Architecture

    Supports: The seven-step pipeline, top-k 8 and the 0.32 gate

  5. Synthesis prompt

    Code

    Supports: Extractive only; the forced answered | refused result

  6. Citation validator

    Code

    Supports: 0 invented citations, by construction

  7. Threshold calibration

    Evaluation

    Supports: Retrieval threshold 0.45 → 0.32

  8. UX critique

    Evaluation31 Aug 2026

    Supports: 22/40 with one P0, “Flagship starter refuses”

  9. Test run

    Test run15 Sep 2026

    Supports: 345 passed / 1 failed (a stale expectation) / 2 skipped, 48 files

  10. Live /api/stats (opens in a new tab)

    Live data15 Sep 2026

    Supports: 5,760 documents and 14,406 chunks

  11. Nightly-crawl design

    Design3 Sep 2026

    Supports: “Staleness is … a correctness bug”

  12. Build ledger

    Build ledger

    Supports: A live smoke test caught Sonnet 5 rejecting a temperature parameter