Featured build · AI systems case study

In Development

PokéJudge AI

A source-grounded decision-support system that investigates an incomplete judge call before recommending a ruling.

Assessment

Insufficient — 1 clarifying question needed.

Clarifying question

How many minutes late did the competitor arrive?

Judge

Exactly 7 minutes after the round started.

Sufficient — generating ruling...

Recommendation

Assess a penalty for Major Tardiness.

Validated Source Support: Strong

2 turns · 2 explicit citations · no conflicts

Role
Sole developer & product designer
Status
In development · local .NET console app
Outcome
End-to-end pipeline with grounding validation and Source Support
Stack
C#, .NET, Gemini embeddings, xUnit

Why I built it

Policy should drive the next question

As a Pokémon Professor and tournament judge, I wanted a tool that could find the right policy quickly without guessing at the game state. Judge calls are time-sensitive, often incomplete, and governed by official material spread across large, cross-referenced documents.

The problem

A fluent answer is not enough

A generic chatbot can fill gaps with plausible assumptions. PokéJudge retrieves first so the source material, not a model’s general Pokémon knowledge, determines which unknown facts matter and whether the evidence supports a recommendation.

Decision pipeline

The application owns the workflow

Built

Describe is the entry point. Retrieval happens before clarification; when the assessment finds a material unknown, the answer updates structured state and triggers a new search before assessment continues.

  1. Step 1

    Describe

    Capture the judge call in the judge’s own words.

  2. Step 2

    Retrieve

    Search authoritative policy before deciding what is missing.

  3. Step 3

    Assess

    Separate sufficient facts from material unknowns.

  4. Step 4

    Clarify

    Ask only questions tied to the retrieved passages.

  5. Step 5

    Re-retrieve

    Search again with newly confirmed facts.

  6. Step 6

    Recommend

    Generate guidance only when the facts are sufficient.

  7. Step 7

    Validate

    Assign Strong, Partial, or Insufficient Source Support.

Clarification loop ↶

Clarify → re-retrieve → assess repeats until the facts are sufficient. Only then does the path continue to Recommend and Validate.

Live product evidence

The complete clarification run

Open full-size screenshot (opens in a new tab)

This capture preserves the full output from the live Gemini-backed late-arrival scenario: both retrieval passes, the material question and answer, the recommendation, validated Source Support, and citation checks.

Captured from a live run using the official-policy corpusComplete terminal output

Complete PokéJudge clarification run

Complete PokeJudge console output for a late-arrival scenario, including retrieval, a clarifying question, re-retrieval, a Major Tardiness recommendation, and validated Strong Source Support.

Captured from a live run using the official-policy corpus · Complete terminal output

System structure

Seven focused components, one controlled path

Ingestion + chunking

Turns official PDFs into normalized, section-aware, citable chunks.

Why it exists: Preserves the authority and location behind every passage.

Embeddings + retrieval

Finds policy passages relevant to the evolving scenario.

Why it exists: Keeps the investigation grounded in the ingested corpus.

Structured scenario state

Tracks confirmed facts, unknowns, and hypotheses separately.

Why it exists: Prevents an interpretation from silently becoming evidence.

Sufficiency + clarification

Decides whether material facts are missing and formulates targeted questions.

Why it exists: Blocks premature rulings.

Ruling generation

Produces a structured recommendation, explanation, repair steps, and citations.

Why it exists: Constrains output to the decision the workflow has earned.

Grounding + Source Support

Checks citation existence, coverage, sufficiency, and conflicts.

Why it exists: Reports evidentiary support instead of a persuasive confidence score.

Evaluation harness

Scores retrieval, clarification, ruling, and grounding across repeatable scenarios.

Why it exists: Makes failures in the path visible, even when the final answer sounds right.

“Confirmed facts, unknown facts, and possible interpretations are different states. A hypothesis can guide the next question, but it can never support a ruling.”

Engineering decisions

Measure support and score the path

Source Support, not model confidence

Strong, Partial, or Insufficient is derived from retrieved authority, citation coverage, fact sufficiency, and source conflict. It describes the available evidence, not how persuasive the model sounds.

Evaluation includes the investigation

The harness scores clarification, retrieval, ruling, and grounding. A correct-looking answer reached through the wrong path does not count as full success.

A source-coverage edge case

Clarification can reveal that the available rule stops short

In a live extra-card run, the system asked whether the card was identifiable and whether the player had seen it, then re-retrieved after each answer. The known facts ultimately fell outside the repair described by the retrieved passage, so the model reported Insufficient support for a concrete remedy. The validator rated the cited passage Strong only for the narrower conclusion that this repair did not apply, not for a penalty the corpus could not support.

This is useful backlog evidence: expand authorized source coverage and make the distinction between “strongly supported limitation” and “strongly supported ruling” clearer in the product language.

Current evidence

Useful regression signals, carefully scoped

20
hand-authored scenarios
Across major judge-call categories and incomplete prompts.
214
deterministic tests
Passing at the Milestone 8.5 implementation checkpoint.
Repeated
live runs
Separating reasoning, retrieval, source-coverage, and provider failures.

This supports regression testing, not a production accuracy claim.

Current state

A complete local workflow

The console application exercises the full pipeline and evaluation harness. There is no web UI yet.

Built
  • End-to-end local console pipeline
  • Official PDF ingestion and section-aware chunking
  • Gemini embeddings, retrieval, and structured model output
  • Clarification loop with structured scenario state
  • Grounding validation and Source Support assignment
  • Deterministic tests and a scenario evaluation harness

What’s next

Harden first, then add a web surface

The next work focuses on consistency, retrieval quality, authorized source coverage, and stronger evaluation before adding scenario entry, short clarifications, visible known facts, Source Support, and expandable citations to a web experience.

Planned next
  • Make repeated-run clarification behavior more consistent
  • Improve ranking within the existing official-policy corpus
  • Add authorized, separately ingested card-specific rulings
  • Expand evaluation and close documented source gaps
  • Build a React/TypeScript and ASP.NET Core web surface after hardening

Explore the work

Grounded by design, honest about the gaps

Next in the set

Featured build

In Development
Loot Singles order detail screen showing sample-order cards and their set, condition, variant, and quantity details.

Loot Singles Fulfillment

A set-aware picking app built to prevent wrong-card mistakes and order collisions.

  • React
  • TypeScript
  • C#

Case study 2 of 5. Up next: Loot Singles Fulfillment.

Featured build

V1 Complete
RoleSync membership tiers listing Silver, Gold and Bronze, each with its Discord role, Shopify customer tag and priority.

RoleSync

A Shopify app that ties member discounts to verified Discord roles.

  • TypeScript
  • React Router
  • Shopify API

Let's connect

If you're hiring, have a project in mind, or want to talk shop, I'd love to hear from you.

Email Me