_ _ _ _ ____ _ _
| | | | __ _ ___| |__ ___ __ _| |_| _ \ ___ ___ ___| |_| |_ __ _
| |_| |/ _` / __| '_ \ / __/ _` | __| |_) / _ \/ __|/ _ \ __| __/ _` |
| _ | (_| \__ \ | | | (_| (_| | |_| _ < (_) \__ \ __/ |_| || (_| |
|_| |_|\__,_|___/_| |_|\___\__,_|\__|_| \_\___/|___/\___|\__|\__\__,_|
Decode the Rosetta Stone of Password Cracking Rules
A Python project designed to analyze hashcat debug mode 4 and mode 5 output files to identify the most efficient rules and track baseword frequency patterns used during password cracking attacks.
- Parse hashcat debug files (
--debug-mode 4and--debug-mode 5) with automatic baseword and rule extraction - Attribute candidates to source wordlists (mode 5) for per-wordlist statistics
- Track rule efficiency by multiple metrics:
- Application frequency (most commonly applied rules)
- Baseword spread (rules applied to most unique basewords)
- Candidate generation (rules producing most unique candidates)
- Monitor baseword patterns with detailed occurrence logs and statistics
- Generate detailed reports with rule and baseword analytics
- Export analysis to JSON or CSV formats for further processing
- Command-line interface for easy analysis and reporting
If you're using uv, you can run without installation:
# Clone the repository
git clone https://github.com/bandrel/HashcatRosetta.git
cd HashcatRosetta
# Run as a module (recommended)
uv run python -m hashcat_rosetta --help
# Or use the installed command
uv run hashcat-rosetta --helpuv tool install git+https://github.com/bandrel/HashcatRosetta.gitThe dev tools live in the dev dependency group.
# With uv (installs the dev group by default)
uv sync
# With pip (25.1+)
pip install -e . --group devAnalyze a hashcat debug file (shows summary by default):
hashcat-rosetta debug_output.txtShow top rules by frequency:
hashcat-rosetta debug_output.txt --rules --top 10 --metric frequencyShow top rules by other metrics:
hashcat-rosetta debug_output.txt --rules --metric basewords
hashcat-rosetta debug_output.txt --rules --metric candidatesShow basewords appearing multiple times:
hashcat-rosetta debug_output.txt --basewords --top 10Show top wordlists (debug mode 5 only):
hashcat-rosetta debug_output.txt --wordlists --top 10Show detailed per-wordlist statistics (unique basewords, candidates, and rules):
hashcat-rosetta debug_output.txt --wordlists --top 10 --detailThe --wordlists output mirrors --rules: a Top N Wordlists header followed by
numbered Wordlist: <name> (<count>) lines. When a mode-5 file is analyzed without
any output flags, the default summary also includes a Wordlist Statistics section.
Force a specific debug mode instead of auto-detecting it:
hashcat-rosetta debug_output.txt --debug-mode 5 --wordlistsShow detailed baseword analysis:
hashcat-rosetta debug_output.txt --basewords --top 10 --detail --min-occurrences 2Export complete analysis report:
hashcat-rosetta debug_output.txt --export report.json --format json
hashcat-rosetta debug_output.txt --export report.csv --format csvExplain what a hashcat rule does step-by-step:
hashcat-rosetta --explain "c$1" --baseword admin
hashcat-rosetta --explain "u$!" --baseword mywordGenerate hashcat masks from English descriptions using a local LLM:
hashcat-rosetta --mask "The word 'Summer' followed by six digits."Output:
Mask Suggestions for: 'The word 'Summer' followed by six digits.'
======================================================================
1. Summer?d?d?d?d?d?d
literal "Summer", then 6 × digit → 1,000,000 candidates
Why: matches the literal word followed by a 6-digit number
Save the generated mask to a file:
hashcat-rosetta --mask "The word 'Summer' followed by six digits." -o masks.hcmaskGenerate masks from other descriptions:
hashcat-rosetta --mask "a capitalized season, two digits, and a special char"
hashcat-rosetta --mask "year 2020-2025 followed by exclamation or question mark"The mask generation feature uses a local Ollama server running an OpenAI-compatible chat
endpoint. By default, it connects to http://localhost:11434 and uses the model
gemma3:27b (see below). These can be configured via environment variables
or CLI flags:
# Using environment variables
OLLAMA_HOST=http://192.168.1.100:11434 OLLAMA_MODEL=llama2:70b \
hashcat-rosetta --mask "your description here"
# Using CLI flags (override environment variables)
hashcat-rosetta --mask "your description" --ollama-host http://custom.host:11434 --model llama2Security note: Mask descriptions are sent only to the Ollama endpoint you configure
(localhost by default, or wherever --ollama-host/OLLAMA_HOST points) — never to a
cloud provider. The OpenAI SDK is used purely as an HTTP client against that endpoint;
no data or API key is ever transmitted to api.openai.com.
The default is chosen by scripts/benchmark_mask_models.py, which runs a fixed set of
14 --mask-style prompts — including custom-charset back-references and category-recall
prompts (Bible books, Bible verse references, European capital cities) — against every
locally-installed candidate model and grades each response three ways:
- Deterministic gate. Each prompt has a hand-written checker (correct keyspace, correct token counts, no duplicate suggestions, etc). A hard fail means the model's response couldn't even be parsed into a valid mask (bad JSON, an unresponsive server); a soft fail means it parsed fine but didn't satisfy the request (wrong digit count, non-vowel custom charset, ...). Soft-failed prompts still get judged, not silently excluded.
- Keyspace cross-check. Every validated suggestion's keyspace is independently
verified against hashcat-utils'
mp64(maskprocessor), when installed — this is the same checkgenerate_masks()itself runs on every suggestion in production, so a benchmark hard fail here also means real--maskusage would have rejected it. - LLM judge. A model on a separate host (
gemma3:12bon a second machine, chosen specifically because it isn't itself a candidate — avoids self-grading bias) scores every response 1-5 for how well it satisfies the original request. Requests are sent with thinking enabled, since a slow model already costs the round-trip time either way.
The recommended default is the smallest model with zero hard fails and a mean judge score ≥ 4. The three finalists carried forward from earlier rounds, re-run against the full 14-prompt set:
| model | size | hard fails | mean score | time |
|---|---|---|---|---|
gemma3:27b |
16.2 GB | 0 | 4.1 | 180s |
dengcao/Qwen3-30B-A3B-Instruct-2507:latest |
17.4 GB | 1 | 4.5 | 176s |
laguna-xs-2.1:latest |
18.9 GB | 2 | 4.7 | 730s |
gemma3:27b is the only one of the three with zero hard fails, so it's the pick despite
not having the highest raw score — dengcao and laguna-xs-2.1:latest scored higher but
each failed at least one prompt outright (and laguna-xs-2.1:latest is also far slower,
730s vs ~180s, driven by its own very large native context window).
Several Qwen3-family models in earlier sweep rounds (qwen3:8b/30b/32b, qwen3.5:9b/27b)
showed a distinct failure mode: hidden "thinking" tokens plus a huge native context window
causing multi-minute-to-30-minute hangs on --mask-scale hardware. The prior default,
qwen3.6:35b-a3b, was replaced for the same reason (see CHANGELOG.md).
Re-run the sweep yourself with uv run python scripts/benchmark_mask_models.py — it pulls
any missing candidates and prints an updated recommendation.
from hashcat_rosetta import DebugAnalyzer
analyzer = DebugAnalyzer()
# Analyze a debug file
result = analyzer.analyze_debug_file('debug_output.txt')
print(f"Total entries: {result['total_entries']}")
print(f"Unique rules: {result['unique_rules']}")
print(f"Unique basewords: {result['unique_basewords']}")
# Get top rules by frequency
top_rules = analyzer.get_top_rules_by_frequency(10)
for rule, count in top_rules:
print(f"Rule: {rule}, Applications: {count}")
# Get top basewords
top_basewords = analyzer.get_top_basewords_by_frequency(10)
for baseword, count in top_basewords:
print(f"Baseword: {baseword}, Occurrences: {count}")
# Get basewords appearing multiple times
frequent_basewords = analyzer.get_basewords_with_min_occurrences(2)
print(f"Basewords appearing 2+ times: {len(frequent_basewords)}")
# Get detailed information about a specific baseword
detail = analyzer.get_baseword_detail('password')
print(f"Rules applied to 'password': {detail['unique_rules']}")
print(f"Occurrences: {len(detail['occurrences'])}")
# Export complete analysis
export = analyzer.export_to_dict()The analyzer automatically detects and supports both hashcat debug output formats:
baseword:rule:candidate
COMPUTER:} } } } t:retupmoc
EXAMPLE:sa@ se3 so0:3x@mpl3
admin:$1 $5 c ^@:@Admin15
Each line contains three colon-separated fields:
- baseword: The original dictionary word
- rule: The hashcat rule applied
- candidate: The resulting password candidate after applying the rule
hashcat has always emitted this format (src/debugfile.c writes orig, :, rule, :, mod).
baseword rule candidate
password c P@ssword
password u PASSWORD
admin l admin
letmein [ etmein
Each line contains three space-separated fields with the same meaning as above. This is an older, legacy format that this parser also accepts.
Note: The analyzer automatically detects which format your file uses. No manual configuration needed!
Hashcat --debug-mode 5 adds a trailing wordlist field to each colon-separated line:
baseword:rule:candidate:wordlist
password:c:P@ssword:/opt/wordlists/rockyou.txt
admin:l:admin:/opt/wordlists/rockyou.txt
letmein:[:etmein:<stdin>
Each line contains four fields:
- baseword: The original dictionary word
- rule: The hashcat rule applied
- candidate: The resulting password candidate
- wordlist: The source dictionary path, or a sentinel (
<stdin>,<generic>,<none>) when hashcat has no path to report
Mode 5 unlocks the --wordlists output and a Wordlist Statistics section in the
default summary. Mode-4 analysis is unchanged.
By default the analyzer auto-detects the mode by counting fields. Use --debug-mode
to force interpretation:
hashcat-rosetta debug.txt --debug-mode auto # default: detect from field count
hashcat-rosetta debug.txt --debug-mode 4 # force mode 4
hashcat-rosetta debug.txt --debug-mode 5 # force mode 5--debug-mode applies to debug-file analysis only (not --analyze-rules).
Windows path limitation: Hashcat does not escape colons, so basewords, candidates,
and Windows wordlist paths (e.g. C:\wordlists\rockyou.txt) may contain :. The parser
assumes the trailing wordlist field contains no colon, which holds for Linux paths and the
sentinels but not for Windows drive-letter paths. Forcing --debug-mode 5 mitigates this
by treating everything after the candidate as the wordlist field.
Generate debug output with hashcat using --debug-mode 4 or --debug-mode 5:
# Mode 4 (baseword rule candidate)
hashcat -m [hash-mode] -a 0 --debug-mode 4 -r rules.rule hashes.txt wordlist.txt > debug.txt
# Mode 5 (baseword:rule:candidate:wordlist — adds source wordlist attribution)
hashcat -m [hash-mode] -a 0 --debug-mode 5 -r rules.rule hashes.txt wordlist.txt > debug.txtImportant: Use --debug-mode 4 or --debug-mode 5. Other debug modes (1-3) produce different output formats that are not compatible with this analyzer.
HashcatRosetta/
├── hashcat_rosetta/ # Main package
│ ├── __init__.py # Package initialization and public API
│ ├── __main__.py # Module entry point (python -m hashcat_rosetta)
│ ├── parser.py # Rule and debug log parsing
│ ├── analyzer.py # Static rule analysis logic
│ ├── debug_analyzer.py # Debug file analysis logic
│ ├── formatting.py # Rule opcode descriptions and display
│ └── cli.py # Command-line interface
├── tests/ # Test suite
│ ├── test_analyzer.py # Tests for analyzers and parsers
│ ├── test_cli.py # CLI interface tests
│ ├── test_edge_cases.py # Edge case and regression tests
│ ├── test_fixes.py # Bug fix verification tests
│ └── test_rule_matrix.py # Rule matrix tests
├── pyproject.toml # Python packaging and tool config
├── LICENSE # MIT license
└── README.md # This file
pytestpytest --cov=hashcat_rosetta tests/ruff check hashcat_rosetta/ tests/
ruff format hashcat_rosetta/ tests/Frequency: How many times a rule was applied across the debug file. Rules with high frequency are the most commonly used.
Unique Basewords: How many different basewords a rule was applied to. Rules affecting more diverse basewords may have broader applicability.
Unique Candidates: How many unique candidate passwords a rule generated. Rules generating more unique candidates may be more valuable for password cracking.
The analyzer tracks every occurrence of each baseword, including:
- Which rules were applied
- What candidates were generated
- The order of operations
This helps identify:
- Most frequently used dictionary words
- Which rules are most effective on specific basewords
- Patterns in word transformation across your attack
hashcat-rosetta FILE Show analysis summary
hashcat-rosetta FILE --rules --metric frequency Show top rules by metric
hashcat-rosetta FILE --basewords --detail Show baseword analysis
hashcat-rosetta FILE --wordlists --detail Show wordlist analysis (mode 5)
hashcat-rosetta FILE --debug-mode 5 --wordlists Force mode 5, show wordlists
hashcat-rosetta FILE --export report.json Export analysis report
hashcat-rosetta --explain "c$1" --baseword admin Explain a rule step-by-step
hashcat-rosetta rules.txt --analyze-rules Analyze rule file opcodes
hashcat-rosetta debug.txt --basewords --min-occurrences 5analyzer = DebugAnalyzer()
analyzer.analyze_debug_file('debug.txt')
detail = analyzer.get_baseword_detail('password')
print(f"Baseword: {detail['baseword']}")
print(f"Occurrences: {detail['total_occurrences']}")
print(f"Rules used: {detail['unique_rules']}")
print(f"Candidates: {detail['unique_candidates']}")
for occurrence in detail['occurrences']:
print(f" Rule: {occurrence['rule']} -> {occurrence['candidate']}")rule_stats = analyzer.get_rule_statistics_summary()
baseword_stats = analyzer.get_baseword_statistics_summary()
print(f"Total rules: {rule_stats['total_rules']}")
print(f"Average applications: {rule_stats['avg_applications_per_rule']:.2f}")
print(f"Total basewords: {baseword_stats['total_basewords']}")Do not use: uv run hashcat_rosetta (this won't work)
Use instead:
# Method 1: Use the CLI command name
uv run hashcat-rosetta debug.txt
# Method 2: Run as a Python module
uv run python -m hashcat_rosetta debug.txtWhy? hashcat_rosetta is the Python package name (for imports), while hashcat-rosetta is the CLI command name (for running). The package name hashcat_rosetta is not executable on its own.
If you see ImportError: attempted relative import with no known parent package, make sure you're running the package as a module with -m:
python -m hashcat_rosetta # Correct
python hashcat_rosetta # Won't workEdit pyproject.toml to customize:
- Project version
- Dependencies
- Development tool configurations
- Entry points
- Create a feature branch
- Make your changes
- Write/update tests
- Run tests and linting
- Submit a pull request
See CHANGELOG.md. It is the single record of what changed in each release; this section used to duplicate it, which is a good way to end up with two histories that disagree.
MIT License - see LICENSE file for details