I build AI systems that prove their answers — production-grade, solo, in days.

Deterministic engines, agent tooling, and evaluation infrastructure. Everything below is shipped, public, and checkable — dates and numbers included.

Shipped, verifiable work

agent-vigil — deterministic trust reports for AI agent sessions

Open-source verifier: extracts every claim an agent makes and checks each against repo reality — tests rerun, files diffed, stuck-loops flagged. 5 detectors, no LLM in the verification path. First bug it caught was in its own author. github.com/sulmusic2-star/agent-vigil

GroundTruth-Geo — open evaluation benchmark for grounded geo answers

Benchmark + deterministic grader + MCP server, MIT-licensed, 8 topic domains, oracle-backed scoring with closed-book frontier-model baseline runs. Built and published solo. github.com/sulmusic2-star/groundtruth-geo

Lasting Ground — a cited-answer engine for address-level questions

Deterministic pipeline that answers property questions with the government source attached to every line: 29,672 permit records across 136 Massachusetts towns, statewide-parcel flood analysis, live national API. No LLM in the answer path — every value traceable. lastingground.com

Disaster-records analysis at claim scale

Post-fire reconstruction of unrecorded structures across 3,131 destroyed parcels by joining decades of permit archives against assessor rolls — producing parcel-level, dollar-quantified findings with a citation per row, live as public evidence pages.

Deterministic reconciliation engine — built and tested in one day

Shopify↔QuickBooks audit engine: 4 error-class detectors (unbooked payouts, zero-fee postings, FX mismatches, duplicate revenue), full test suite, CSV-ingest path, working product UI — idea to verified build in a single day.

Agent-first operating practice

Daily production use of Claude Code, multi-agent research fleets with adversarial verification stages, MCP servers, and deterministic scoring harnesses — the workflow itself is part of the portfolio.

What I'm best at

Verification-grade data engineering

Systems where every output carries its evidence — graders, reconciliation, provenance, eval sets.

Speed

Production v1s in days, not quarters — solo, end to end: engine, tests, UI, deploy.

Tim Sullivan · [email protected] · 781-363-0505