Does increasing GPT-5.2 reasoning effort improve diagnosis accuracy enough to justify the token/latency cost? Ablation study on 897 paired medical cases.
-
Updated
Jul 9, 2026 - Python
Does increasing GPT-5.2 reasoning effort improve diagnosis accuracy enough to justify the token/latency cost? Ablation study on 897 paired medical cases.
可复核的 Agent 评测工具链:Manifest、失败归因、McNemar、Bootstrap 与公开证据边界。
Measurement trust + trace-to-training backend — eval_trust audit toolkit and rollout→SFT/DPO/RL data factory for WasmAgent compliance training
Classification models for detecting fake reviews and predicting software bugs. Includes implementations of decision trees, bagging, random forests, logistic regression, and Naive Bayes, with statistical evaluation using McNemar's test.
How much of a SWE-bench leaderboard gap is real? On Verified, 132 of 133 adjacent gaps are below the benchmark's own detection floor and none survives a family-wise correction.
A hybrid anomaly detection pipeline combining ensemble machine learning models and deep learning techniques for credit card fraud detection. Evaluates 25 model combinations across multiple datasets and validates performance using McNemar and Friedman statistical tests.
Paste two eval runs, or just the counts, and find out whether that 3-point score move is a real improvement or sampling noise.
Statistical tooling for LLM evaluation decisions: paired comparison, judge stability, sample sizing, and a ship/hold verdict.
Pre-registered cross-validated voter committees for honest evaluation on dental panoramic VQA: 75.36% on MMOral-OPG-Bench (370/491), McNemar p=1.1e-7
Add a description, image, and links to the mcnemar-test topic page so that developers can more easily learn about it.
To associate your repository with the mcnemar-test topic, visit your repo's landing page and select "manage topics."