I build local-first, deterministic, auditable tools for AI agents and developer workflows.
我在台灣打造本機優先、可重播、可稽核的 AI agent 與開發工具。
Models may propose. Verifiers decide. Missing evidence stays
UNKNOWN/INCOMPLETE.
Frontier Atlas is a private product and research line for auditing whether an agent's claim is actually supported by its cited evidence.
frontier-atlas-open-tests
is its public, offline test surface:
- bounded claim–citation audit schemas;
- deterministic protocol and identity verification;
- blind, non-gold calibration packets;
- commit → reveal → adjudication workflows;
- reproducible public issue intake.
The current release stage is P4.5-T: external tester preparation. Promotion still requires a second, independent natural person; that rule is not replaced by two models, two sessions, or two signing keys.
The planned path is:
external qualification → double-blind pilot + real-source calibration →
private judge/token optimization → development benchmark →
sealed holdout → black-box closed beta → release candidate →
external evidence review → production
Public visibility grants neither a passing result nor access to private product core, hidden labels, qualification keys, human records, or sealed holdout authority. A repository is open source only when its own license explicitly says so.
| Project | Purpose |
|---|---|
| frontier-atlas-open-tests | Public offline qualification and semantic-audit test surface |
| greenwash | Detects when an agent makes CI green by weakening verification |
| nullbench | Pre-registers decisions and scores them against chance |
| trust-meter | Deterministic, evidence-backed trust scoring |
| unasked | Evidence-gated repository investigation; non-certifying |
| branchback | Replays belief-at-the-time against knowledge-now |
| aurora | Evidence-led industry discovery without an LLM at runtime |
| receiptradar | Local receipt-to-ledger CLI with no cloud account |
| universe-explorer | Epistemically honest science knowledge system |
- Evidence over confidence. Claims point to exact citations, tests, diffs, logs, artifacts, or an explicit insufficient-data result.
- Fail closed. Missing observation is not a pass.
- Deterministic first. The same frozen evidence should produce the same mechanical result.
- Authority stays separate. A model can propose; it cannot silently grant review, qualification, or release authority to itself.
- Keep negative results. Blocked gates and counterexamples remain visible rather than being rewritten as success.
- Measure cost with real denominators. Speed, token, and accuracy claims wait for reproducible benchmark data.
Python · Rust · Go · TypeScript · FastAPI · React ·
SQLite · offline-first CLI workflows
Public portfolio reconciled 2026-08-29. Current status: preparing Frontier Atlas for public semantic-audit testing.