Ingestion + chunking
Turns official PDFs into normalized, section-aware, citable chunks.
Why it exists: Preserves the authority and location behind every passage.
Featured build · AI systems case study
In DevelopmentA source-grounded decision-support system that investigates an incomplete judge call before recommending a ruling.
Assessment
Insufficient — 1 clarifying question needed.
Clarifying question
How many minutes late did the competitor arrive?
Judge
Exactly 7 minutes after the round started.
Sufficient — generating ruling...
Recommendation
Assess a penalty for Major Tardiness.
Validated Source Support: Strong
2 turns · 2 explicit citations · no conflicts
Why I built it
As a Pokémon Professor and tournament judge, I wanted a tool that could find the right policy quickly without guessing at the game state. Judge calls are time-sensitive, often incomplete, and governed by official material spread across large, cross-referenced documents.
The problem
A generic chatbot can fill gaps with plausible assumptions. PokéJudge retrieves first so the source material, not a model’s general Pokémon knowledge, determines which unknown facts matter and whether the evidence supports a recommendation.
Decision pipeline
Describe is the entry point. Retrieval happens before clarification; when the assessment finds a material unknown, the answer updates structured state and triggers a new search before assessment continues.
Step 1
Capture the judge call in the judge’s own words.
Step 2
Search authoritative policy before deciding what is missing.
Step 3
Separate sufficient facts from material unknowns.
Step 4
Ask only questions tied to the retrieved passages.
Step 5
Search again with newly confirmed facts.
Step 6
Generate guidance only when the facts are sufficient.
Step 7
Assign Strong, Partial, or Insufficient Source Support.
Clarification loop ↶
Clarify → re-retrieve → assess repeats until the facts are sufficient. Only then does the path continue to Recommend and Validate.
Live product evidence
This capture preserves the full output from the live Gemini-backed late-arrival scenario: both retrieval passes, the material question and answer, the recommendation, validated Source Support, and citation checks.
System structure
Turns official PDFs into normalized, section-aware, citable chunks.
Why it exists: Preserves the authority and location behind every passage.
Finds policy passages relevant to the evolving scenario.
Why it exists: Keeps the investigation grounded in the ingested corpus.
Tracks confirmed facts, unknowns, and hypotheses separately.
Why it exists: Prevents an interpretation from silently becoming evidence.
Decides whether material facts are missing and formulates targeted questions.
Why it exists: Blocks premature rulings.
Produces a structured recommendation, explanation, repair steps, and citations.
Why it exists: Constrains output to the decision the workflow has earned.
Checks citation existence, coverage, sufficiency, and conflicts.
Why it exists: Reports evidentiary support instead of a persuasive confidence score.
Scores retrieval, clarification, ruling, and grounding across repeatable scenarios.
Why it exists: Makes failures in the path visible, even when the final answer sounds right.
“Confirmed facts, unknown facts, and possible interpretations are different states. A hypothesis can guide the next question, but it can never support a ruling.”
Engineering decisions
Strong, Partial, or Insufficient is derived from retrieved authority, citation coverage, fact sufficiency, and source conflict. It describes the available evidence, not how persuasive the model sounds.
The harness scores clarification, retrieval, ruling, and grounding. A correct-looking answer reached through the wrong path does not count as full success.
A source-coverage edge case
In a live extra-card run, the system asked whether the card was identifiable and whether the player had seen it, then re-retrieved after each answer. The known facts ultimately fell outside the repair described by the retrieved passage, so the model reported Insufficient support for a concrete remedy. The validator rated the cited passage Strong only for the narrower conclusion that this repair did not apply, not for a penalty the corpus could not support.
This is useful backlog evidence: expand authorized source coverage and make the distinction between “strongly supported limitation” and “strongly supported ruling” clearer in the product language.
Current evidence
This supports regression testing, not a production accuracy claim.
Current state
The console application exercises the full pipeline and evaluation harness. There is no web UI yet.
What’s next
The next work focuses on consistency, retrieval quality, authorized source coverage, and stronger evaluation before adding scenario entry, short clarifications, visible known facts, Source Support, and expandable citations to a web experience.
Explore the work