Skip to content

Upgrade to C++20, update lc0 integration, and add one-game-per-file output - #1

Draft
ContradNamiseb wants to merge 56 commits into
CallOn84:masterfrom
Bonan14:master
Draft

Upgrade to C++20, update lc0 integration, and add one-game-per-file output#1
ContradNamiseb wants to merge 56 commits into
CallOn84:masterfrom
Bonan14:master

Conversation

@ContradNamiseb

@ContradNamiseb ContradNamiseb commented Jan 7, 2026

Copy link
Copy Markdown

This pull request modernizes the build system and codebase, updates the lc0 integration, and introduces several improvements to training data handling and developer tooling. The main changes include upgrading to C++20, updating the lc0 submodule and source file paths, switching to the V6 training data format, and enhancing the output format to write one game per chunk file. Additionally, new calibration scripts and an absl compatibility header were added.

Build System and Language Modernization

  • Upgraded the project to use C++20 (from C++17), set explicit Release build type, simplified include directories, and added lc0/src/utils/string.cc to resolve linker errors. The build now prefers the system zlib on Unix and includes platform-specific compiler flags for better compatibility and legacy code support. (CMakeLists.txt, .github/workflows/cmake-multi-platform.yml) [1] [2] [3]

lc0 Integration Updates

  • Updated the lc0 submodule to track the official master branch and adjusted source file paths to match the latest lc0 structure (e.g., removed bitboard.cc, replaced neural/writer.cc with trainingdata/writer.cc). (.gitmodules, lc0) [1] [2]

Training Data Format and Output Improvements

  • Upgraded from V4 to V6 training data format, updated related source files and hash utilities, and modified the output logic so each game is written to its own chunk file instead of batching multiple games together. (PULL_REQUEST.md)

Developer Tooling and Scripts

  • Added three new Python scripts (calibrate_s.py, calibrate_s2.py, calibrate_s3.py) for analyzing and calibrating the relationship between Q and D values in training data, supporting improved model calibration and validation. (scripts/calibrate_s.py, scripts/calibrate_s2.py, scripts/calibrate_s3.py) [1] [2] [3]

Dependency Compatibility

  • Introduced a minimal absl/cleanup/cleanup.h header to satisfy new lc0 dependencies. (absl/cleanup/cleanup.h)## Summary
    This PR modernizes the trainingdata-tool by upgrading to C++20, updating the lc0 integration to work with the latest lc0 codebase, and improving the output format to write one game per chunk file.

Changes

Build System Updates

  • Upgraded C++ standard from C++17 to C++20
  • Set explicit Release build type
  • Simplified include directories
  • Added lc0/src/utils/string.cc to fix linker error for StrSplit function

lc0 Integration Updates

  • Updated lc0 source file paths to match latest lc0 structure:
    • Removed lc0/src/chess/bitboard.cc (no longer needed)
    • Changed lc0/src/neural/writer.cc to lc0/src/trainingdata/writer.cc
  • Updated .gitmodules for lc0 submodule
  • Added absl/ library dependency

Training Data Format Upgrade

  • Upgraded from V4 to V6 training data format
  • Replaced V4TrainingDataHashUtil.h with V6TrainingDataHashUtil.h
  • Updated related source files for V6 compatibility

Output Format Improvement

  • Modified TrainingDataWriter to write one game per chunk file
  • Each game's training positions are now isolated in their own .gz file
  • Removed batching logic that previously combined multiple games into one file

Testing

  • Successfully built on Linux with GCC
  • Verified output: Converting 5 games produces 5 separate files with varying sizes reflecting different game lengths

Made with Gemini and Claude

But dont merge just yet I still have to manually verify this code.

- Upgraded C++ standard from C++17 to C++20
- Updated lc0 source file paths to match new lc0 structure:
  - Removed lc0/src/chess/bitboard.cc (no longer needed)
  - Changed lc0/src/neural/writer.cc to lc0/src/trainingdata/writer.cc
  - Added lc0/src/utils/string.cc (fixes StrSplit linker error)
- Updated include directories (simplified paths, added root include)
- Set explicit Release build type
- Added absl/ library dependency
- Upgraded TrainingData from V4 to V6:
  - Replaced V4TrainingDataHashUtil.h with V6TrainingDataHashUtil.h
  - Updated related source files for V6 compatibility
- Updated .gitmodules for lc0 submodule

Made with Gemini and Claude Opus
- Modified EnqueueChunks() to write all training positions from a single
  game directly to its own .gz file
- Each game now gets its own chunk file (game_XXXXXX.gz)
- Removed the batching logic that combined multiple games into one file
- This ensures training data from different games is not mixed together

Made with Gemini and Claude
@ContradNamiseb
ContradNamiseb marked this pull request as draft January 7, 2026 22:08
- Replaced boost::hash_range and boost::hash_combine with lc0's HashCat()
  from utils/hashcat.h in V6TrainingDataHashUtil.h
- Removed find_package(Boost) and Boost_INCLUDE_DIRS from CMakeLists.txt
- This eliminates the external Boost dependency for easier building

Made with Gemini and Claude
- Added CI/CD workflow for Ubuntu (gcc, clang) and Windows (cl)
- Builds on push to master and pull requests
- Automatically creates pre-release with artifacts on master push
- No Boost dependency required (uses lc0 native HashCat)

Made with Gemini and Claude
- Copied proto/net.pb.h stub from lc0 to src/proto/ (it was untracked
  in lc0 submodule)
- Added 'src' to include_directories so proto/ can be found
- This fixes the 'proto/net.pb.h file not found' error on CI

Made with Gemini and Claude
- Updated CMake to use system zlib (via find_package) on Unix
- Bundled zlib only used on Windows now
- Added zlib1g-dev to CI Linux dependencies
- Added proper project() declaration to fix CMake warnings
- Fixes 'call to undeclared function' errors for lseek/read/write/close

Made with Gemini and Claude
- Added _CRT_SECURE_NO_WARNINGS to suppress deprecation warnings for
  strcpy, sprintf, fopen, etc.
- Added /Zc:strictStrings- to allow const char[] to char* conversion
  (required by polyglot's getopt.h)

Made with Gemini and Claude
- Added /FIarray to force include <array> header (missing in lc0 submodule)
- Added /permissive to relax conformance rules (helps with polyglot's legacy C code)
- Maintained /Zc:strictStrings- for char* conversions

Made with Gemini and Claude
- Moved /Zc:strictStrings- to polyglot-specific file properties
- Added global warning suppressions: /wd4996, /wd4267, /wd4244, /wd4390, /wd4018
- This fixes C2440 errors in polyglot's getopt.h and cleans up the build log

Made with Gemini and Claude
- Switch Windows CI matrix to use gcc/g++ (MinGW) and Ninja generator
- Enable verbose build logging
- Add MinGW compile flags for polyglot (-fpermissive) to fix const char* errors

Made with Gemini and Claude
- Use -iquote for polyglot sources on GCC/MinGW to prevent <getopt.h> from picking up polyglot/src/getopt.h
- This fixes 'undefined reference to getopt_internal' linker error on Windows
- MSVC continues to use standard include path as it requires the local getopt.h polyfill

Made with Gemini and Claude
- Restore polyglot/src to regular include_directories (iquote wasn't working)
- For MinGW: pre-define __GETOPT_H__ to prevent local getopt.h from being
  included when system unistd.h includes <getopt.h>

Made with Gemini and Claude
- Changed artifact upload to only include trainingdata-tool binary
- Fixes 'Failed to upload CMakeCache.txt' error in release step

Made with Gemini and Claude
- Fixed regex patterns: [%eval ...] now correctly escaped as \[%eval ...\]
- Added try-catch around std::stof to prevent crashes on malformed input

Made with Gemini and Claude
- New StockfishEvaluator class with fork/exec for bidirectional UCI protocol
- CLI options: -stockfish <path> and -sf-depth <N>
- Evaluates each position using Stockfish and converts centipawns to Q-value
- Cross-platform: POSIX (Linux/Mac) uses fork/exec, Windows uses popen

Usage: ./trainingdata-tool -stockfish /path/to/sf -sf-depth 15 games.pgn

Made with Gemini and Claude
- Removed outdated Boost dependency
- Added comprehensive options table
- Added Stockfish evaluation examples
- Updated build instructions

Made with Gemini and Claude
- Added select() with 100ms timeout before blocking read()
- Prevents timeout during UCI initialization
- Added sys/select.h include

Made with Gemini and Claude
- Prints 'TrainingData Tool v1.1 (Stockfish Arg Fix)' on startup
- Helps confirm correct binary is running

Made with Gemini and Claude
- Adds case-insensitive check for .pgn extension
- Prevents processing of binary files or other non-PGN inputs
- Adds cross-platform compatibility for strcasecmp

Made with Gemini and Claude
- Prevents tool from hanging if Stockfish doesn't respond
- Returns 0 (draw eval) on timeout with error message
- Critical for robustness with large/problematic PGN files

Made with Gemini and Claude
…lues, and MLH verification

- Implement move-by-move position feeding to Stockfish (position startpos moves ...)
- Add poly_move_to_uci for correct UCI string generation
- Return StockfishResult struct with score_cp, best_move, nodes
- Differentiate root_q (played move eval) from best_q (best move eval)
- Post-process chunks: root_q = -next_chunk.best_q
- Add setPositionMoves to StockfishEvaluator interface
- Update verify_chunks.py with PliesLeft (MLH) reverse increment
- Rename chunks to moves in verification output
- Add robust SAN cleaning (whitespace, move numbers, comments, annotations)
- Add poly_move_to_uci() function for UCI notation conversion
- Change input format to INPUT_CLASSICAL_112_PLANE
- Improve result parsing with strcmp and fallback for unrecognized results
- Graceful illegal move handling (continue instead of break)
- Add best_move and visits parameters to get_v6_training_data
- Store Q values directly (relative to side-to-move)
- Implement StaticEvaluator with material balance (P=100, N=320, B=330, R=500, Q=900)
- Add piece-square tables for all piece types with MG/EG king tables
- Implement pawn structure analysis (doubled, isolated, passed pawns)
- Add simplified mobility scoring
- Use tapered evaluation (MG/EG interpolation based on game phase)
- Integrate into PGNGame for non-Lichess mode
- Use sigmoid function for cp-to-Q conversion
- Implement StaticEvaluator with material balance (P=100, N=320, B=330, R=500, Q=900)
- Add piece-square tables for all piece types with MG/EG king tables
- Implement pawn structure analysis (doubled, isolated, passed pawns)
- Add simplified mobility scoring
- Use tapered evaluation (MG/EG interpolation based on game phase)
- Integrate into PGNGame for non-Lichess mode
- Use sigmoid function for cp-to-Q conversion
…NNIndex

- Validate all MoveToNNIndex results before indexing probabilities array
- Check legal moves loop indices
- Check played_idx before writing probabilities[played_idx] = 1.0
- Check best_idx before storing in result.best_idx
- Fallback to 0 or played_idx when index >= 1858
- Add debug warning for invalid indices

Fixes crash on Windows where MSVC uninitialized memory (0xCCCC = 52428)
exceeds array bounds (1858 elements).
…NNIndex

- Validate all MoveToNNIndex results before indexing probabilities array
- Check legal moves loop indices
- Check played_idx before writing probabilities[played_idx] = 1.0
- Check best_idx before storing in result.best_idx
- Fallback to 0 or played_idx when index >= 1858
- Add debug warning for invalid indices

Fixes crash on Windows where MSVC uninitialized memory (0xCCCC = 52428)
exceeds array bounds (1858 elements).
Root cause of Windows crash: move_from() and move_to() return polyglot
0x88 format square indices, but lczero::Square::FromIdx() expects 0-63.

- Added square_to_64() conversion before creating lczero::Square
- This was causing garbage indices (0xCCCC) and assertion failures
- Fixes: move 'e4' now correctly outputs as 'e2e4' (was showing 'b8a4')
Root cause of Windows crash: move_from() and move_to() return polyglot
0x88 format square indices, but lczero::Square::FromIdx() expects 0-63.

- Added square_to_64() conversion before creating lczero::Square
- This was causing garbage indices (0xCCCC) and assertion failures
- Fixes: move 'e4' now correctly outputs as 'e2e4' (was showing 'b8a4')
ContradNamiseb and others added 26 commits January 9, 2026 14:33
Lc0's internal board representation is always from white's perspective.
After ApplyMove(), Position::Mirror() is called to switch perspective.
The move passed to PositionHistory::Append() must use absolute board
squares, not flipped coordinates.

Fixes: assertion failure 'ourpieces.intersects(BitBoard::FromSquare(move.from()))' in board.cc:597
Lc0's internal board representation is always from white's perspective.
After ApplyMove(), Position::Mirror() is called to switch perspective.
The move passed to PositionHistory::Append() must use absolute board
squares, not flipped coordinates.

Fixes: assertion failure 'ourpieces.intersects(BitBoard::FromSquare(move.from()))' in board.cc:597
- Added is_black_move parameter to poly_move_to_lc0_move() function
- Regular moves for black are now flipped to match LC0's white perspective
- Castling moves are left unflipped as they use perspective-independent file representation
- Prevents assertion failure when applying black's moves to LC0 board

Fixes issue where training data tool crashed when processing games with black moves.
Removed the release job from the CI workflow.
Restores the release job that creates pre-release on each push to master.
- Added is_black_move parameter to poly_move_to_lc0_move() function
- Regular moves for black are now flipped to match LC0's white perspective
- Castling moves are left unflipped as they use perspective-independent file representation
- Prevents assertion failure when applying black's moves to LC0 board

Fixes issue where training data tool crashed when processing games with black moves.
Removed the release job from the CI workflow.
Restores the release job that creates pre-release on each push to master.
- Fix bug in poly_move_to_lc0_move where normal moves and en-passant were incorrectly nested inside the castling check block
- Correctly set result_d for draws (1.0) and wins/losses (0.0)
- Initialize all V6 training data fields: played_q, played_d, played_m, root_d, best_d, root_m, best_m, policy_kld
- This ensures compatibility with the lc0 rescorer tool
…aw, CLI validation, timeout handling, Windows pipes, verify failures

Co-authored-by: ContradNamiseb <20321785+ContradNamiseb@users.noreply.github.com>
… redirect SF stderr to /dev/null

Co-authored-by: ContradNamiseb <20321785+ContradNamiseb@users.noreply.github.com>
Co-authored-by: ContradNamiseb <20321785+ContradNamiseb@users.noreply.github.com>
Co-authored-by: ContradNamiseb <20321785+ContradNamiseb@users.noreply.github.com>
Resolve conflicts in src/PGNGame.cpp, src/StaticEvaluator.cpp,
src/trainingdata-tool.cpp, src/trainingdata.cpp, and src/trainingdata.h
by keeping the stockfish-eval branch versions, which already include the
review fixes for everything ported to master (V6 fields with plies_left/D,
illegal-move abort, NAG filtering, best-move fallback, square_from_64
board indexing, and CLI argument validation) plus the Stockfish
evaluation support.

Co-authored-by: ContradNamiseb <20321785+ContradNamiseb@users.noreply.github.com>
Added: Stockfish eval

- and lot of fixes by copilot
Co-authored-by: ContradNamiseb <20321785+ContradNamiseb@users.noreply.github.com>
Co-authored-by: ContradNamiseb <20321785+ContradNamiseb@users.noreply.github.com>
Co-authored-by: ContradNamiseb <20321785+ContradNamiseb@users.noreply.github.com>
…-job

Fix release workflow asset name collisions
Co-authored-by: ContradNamiseb <20321785+ContradNamiseb@users.noreply.github.com>
Co-authored-by: ContradNamiseb <20321785+ContradNamiseb@users.noreply.github.com>
Align README CLI options with the current parser and default evaluation behavior
…nd conversion pipeline tools

- Implement -pgn-eval-mode to extract evaluations directly from PGN move comments without spawning external engines.
- Add ScoreToWDL and WDLRescale in src/WdlConversion.h based on Lc0's logistic model with configurable -wdl-scale and -wdl-spread.
- Add -visit-budget option to generate dynamic policy probability distributions based on Q-score (0.5 + |Q|/2 share on played move).
- Add 50-move rule damping (-r50-damp-start) to static evaluation.
- Add scripts for Syzygy tablebase rescoring (rescore_all.py, rescore_chunks.py) and archive packing (pack_chunks.py).
- Add calibration and evaluation measurement scripts (measure_pgn_wdl.py, measure_pgn_draw_rate.py, calibrate_s*.py).
- Enhance verify_chunks.py with policy distribution decoding and moves-left validation.
- Update documentation with Windows (MSVC & MinGW) and Linux build instructions, sharpness analysis, and parameter guidance.
@ContradNamiseb
ContradNamiseb marked this pull request as ready for review August 18, 2026 11:15
@ContradNamiseb
ContradNamiseb marked this pull request as draft August 19, 2026 05:44
Two defects, both found by a 29-file conversion that reported success
while keeping a fraction of its output.

The writer was constructed inside convert_games, which runs once per
input file, so its counter restarted at game_000000 for every PGN and
each input silently overwrote the last one's chunks. A run that logged
2,002,215 games left 398,750 files on disk. It is now built once in
main and passed in, so the counter spans the whole run.

EnqueueChunks also held its mutex across create_directories, the file
open, the gzip compression and the close, which serialised every worker
onto one core -- measured 2.5 of 8 busy, and no thread count could have
helped. Only the index reservation needs the lock now; compression
happens outside it. The mkdir stays inside deliberately: moving it out
lets a second worker see the directory marked created and write into it
before create_directories returns, which faults immediately.

Verified end to end on the full corpus: 2,002,208 games across 334
directories with the counter climbing across every file boundary, then
rescored and packed to 338 archives. Throughput went from ~91 to ~290
games/sec. The remaining ceiling is the single-threaded PGN producer,
not the writer.

The two pipeline scripts drive that run unattended and are resumable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants