→ golobokov.dev · hugging face · kaggle
I like projects where the hard part is proving what a system actually does: checking for leakage, picking a baseline that is hard to beat, and printing the limitation next to the result, negative results included.
┌─ now ─────────────────────────────────────────────────────────────┐
│ building claude code plugins · datasets · evals │
│ measuring ML evaluation · data provenance · agent reliability │
│ stack python · polars · scikit-learn · lightgbm · shell │
└───────────────────────────────────────────────────────────────────┘
![]() Rooms for coding agents. A tree of rooms, scoped invites across companies, and every
|
![]() Five small Claude Code plugins that share no state: a multi-model council with anonymised peer review, search over 1,679 public APIs, supply-chain vetting and guardrail hooks.
|
![]() A sound when Claude Code asks for permission or finishes a turn, and on nothing else. Under 250 lines of shell, and it plays any sound you want.
|
![]() EV charging for Prague 2030: a demand forecast at 24.03 kWh MAE (38.8% below baseline, the honest re-run), charger-type classification and a capacity-constrained load-shifting replay.
|
![]() 120,981 multilingual crypto-news events with weak annotations, each joined to what Bitcoin did over the next minute to 24 hours, timed by when the story was actually readable.
|
![]() Can a model pick which lines of a Claude Code tool output the agent will need later? Six selectors against simply keeping the head and tail, on real sessions.
|
More public projects
| bb10-whatsapp | a native WhatsApp client for BlackBerry 10 and Android 4.3+, on a hardened Node backend |
The banner and tiles are rendered frame by frame with my own ASCII engine (numpy + Pillow), in GitHub's own page colours. The VoltPlan field is generated, not competition data.







