Skip to content

About

A step-by-step horse racing API tutorial: test whether a betting strategy has ever worked, on real races and the prices that were actually available, using the Horse Racing API as your horse racing data API and nothing but Python

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

How to build a horse racing backtester with Python

A step-by-step horse racing API tutorial: test whether a betting strategy has ever worked, on real races and the prices that were actually available, using the Horse Racing API as your horse racing data API and nothing but Python and requests. By the end you will have a backtester that measures strike rate, level-stakes return and A/E (actual against expected), and a worked comparison of backing the favourite against backing everything else.

You will need: a key from https://apihorseracing.com/account and the documentation open beside you.

In this repository

File What it is
README.md This guide
backtest.py The complete backtester, one command to run
docs/api-guide.md Every endpoint of the API, with the plan that reaches it and code in curl, Python and PHP
image

Why strike rate alone loses money

A strike rate tells you how often something wins. It says nothing about the price, and the price is the whole game: a method that finds a 40% winner at even money loses, and one that finds a 10% winner at 14/1 wins. Two measures fix that:

  • Level-stakes return. Back every selection for one unit. Add up what came back, subtract what went out. Positive means the idea made money at the prices on offer.
  • A/E, actual against expected. Each runner's starting price implies a chance of winning: a horse at 4.0 decimal implies 1 in 4. Add those up across a race and you get more than 100%, often 120% to 140%; the excess is the bookmaker's margin. Divide each runner's implied chance by that total and you have its fair share of the race, and the fair shares of a field add up to exactly one winner. Add up the fair shares of your selections and you have the winners the market expected; divide the winners you actually got by that. Above 1.00, your selections won more often than the market gave them credit for. A/E tells you whether the market misjudges something; the return tells you whether that is enough to beat the margin.

Both need the starting price. That is why a backtest needs a plan with prices, and why it needs history: one day of results tells you nothing.

Which key you need

Free key Archive, $99 Complete, $119
Reaches yesterday, today and ahead January 2017 to yesterday January 2017 to ahead
Starting prices held back, null with a note yes yes
Allowance 50 a day 150,000 a month 400,000 a month

Write and test the code on the free key, then run it on Archive, the plan built for research: the code does not change, only what the key reaches. On a free key the backtester checks the key, explains why there is nothing to judge and stops, without spending your daily lookups. Compare the plans.

The two calls a backtest needs

  1. GET /v1/results/{date} returns the day's settled races, grouped into meetings. Each race carries a race_id.
  2. GET /v1/races/{race_id}/result returns that race's finishing order. Each runner in result carries position, horse, sp and sp_decimal, and runners who did not finish are listed separately in unplaced, so a horse that pulled up is never counted as a runner that finished last.

Run it

pip install -r requirements.txt
python3 backtest.py

That is the whole of it: one command, start to finish.

  1. Your key. The first run asks for it, checks it against /v1/account/usage (which costs nothing against your allowance) and saves it to .ahr_key beside the script, readable only by you. You are not asked again. AHR_KEY in the environment takes precedence if you prefer that, and --forget-key clears the saved one.
  2. Your plan. It shows what the key reaches, its rate and what is left of the month. On a free key it explains why there is nothing to backtest and stops there, before spending any of your 50 lookups.
  3. The range. Pick the last 7, 14, 30 or 90 days with the arrow keys, or type your own dates. --from 2026-08-01 --to 2026-08-31 skips the question.
  4. The fetch. Every day gets a live progress bar with a spinner, a count of races and how much of the minute's allowance is left. Calls are paced from the key's own X-RateLimit-* headers, so it runs just under the limit instead of into it, and a 429 is waited out on screen rather than failing. Each finished day is kept in ./data, so a range already fetched is never fetched again. Ctrl+C stops early and still reports the days it has.
  5. The result. A table of both rules, bars for strike rate and A/E, level-stakes profit by day (by week or month on longer ranges), and a plain-English reading of each.

The date range is walked newest first. If your plan does not reach the oldest days, the run stops at the edge and reports what it could reach.

image

The backtester

This is the part that does the work, exactly as it appears in backtest.py. The rest of the file is the terminal around it: the key, the menu, the pacing and the drawing.

def races_on(api, day, on_race=None):
    """Every settled race on a day, each with its finishing order.

    A settled result never changes, so a finished day is fetched once and kept
    in ./data. Yesterday is not kept: racing abroad may still be settling.
    Returns (races, came_from_cache).
    """
    cached = CACHE / f"{day}.json"
    if cached.exists():
        return json.loads(cached.read_text()), True

    meetings = api.get(f"/results/{day}")
    total = sum(len(m.get("races", [])) for m in meetings)
    races = []
    for meeting in meetings:
        for race in meeting.get("races", []):
            result = api.get(f"/races/{race['race_id']}/result")
            races.append({"race_id": race["race_id"], "course": meeting.get("course"),
                          "runners": result.get("result", []),
                          "unplaced": result.get("unplaced", [])})
            if on_race:
                on_race(len(races), total)

    if day < (date.today() - timedelta(days=1)).isoformat():
        CACHE.mkdir(exist_ok=True)
        tmp = cached.with_suffix(".tmp")
        tmp.write_text(json.dumps(races))
        tmp.replace(cached)
    return races, False


def backtest(races, select):
    """Back every runner `select` picks, one unit each, at starting price."""
    bets = wins = judged = skipped = 0
    returned = expected = 0.0
    for race in races:
        field = race["runners"] + race["unplaced"]
        priced = [r for r in field if r.get("sp_decimal")]
        if not priced or len(priced) < len(field):
            skipped += 1          # a race without every price cannot be judged fairly
            continue
        book = sum(1 / r["sp_decimal"] for r in priced)      # over 1.00: the margin
        judged += 1
        for runner in select(priced):
            bets += 1
            expected += (1 / runner["sp_decimal"]) / book     # its fair chance
            if runner.get("position") == 1:
                wins += 1
                returned += runner["sp_decimal"]
    return {
        "races": judged,
        "skipped": skipped,
        "bets": bets,
        "winners": wins,
        "profit": returned - bets,
        "strike_rate": wins / bets * 100 if bets else None,
        "roi_pct": (returned - bets) / bets * 100 if bets else None,
        "a_over_e": wins / expected if expected else None,
    }


def favourite(field):
    """The shortest-priced runner (the first listed, if two share the price)."""
    shortest = min(r["sp_decimal"] for r in field)
    return [r for r in field if r["sp_decimal"] == shortest][:1]


def not_favourite(field):
    """Every runner except that favourite."""
    fav = favourite(field)[0]
    return [r for r in field if r is not fav]


RULES = [("Favourite", favourite), ("Everything else", not_favourite)]

Four decisions in there are worth copying into any backtest of your own:

  • A race is judged only if every runner has a price. A missing price would make the favourite wrong and the expected winners short, so the race is left out and counted, rather than guessed at.
  • A/E uses fair chances, not raw prices. Measured against the raw prices, the margin alone would push every selection below 1.00 and make everything look mispriced. With the margin divided out, 1.00 genuinely means priced right, and the margin shows up where it belongs, in the return.
  • Casualties stay in the field. unplaced runners are losing bets like any other; leaving out the horse that pulled up would flatter every rule that backed it.
  • Yesterday is fetched but not cached. Racing abroad can still be settling, so only finished days are written to ./data.
image

Reading the result

The result panel gives each rule's bets, winners, strike rate, level-stakes profit, return and A/E, then draws them. Two things to look for:

  • The favourite wins more often and still usually loses at level stakes. A high strike rate with a negative return is the clearest demonstration of why strike rate alone is the wrong measure.
  • A/E above 1.00 is not the same as profit. Favourites typically win more often than their share of the market, an A/E above 1.00, and still lose at level stakes, because the margin is taken out of the price first. A method is only interesting when its return holds up across a long, varied period, not a lucky month.

This guide does not quote figures for you: run it on your own range and read your own numbers. A 30-day window is a smoke test, not evidence. Choose your own dates and widen the range to a full season or more before believing anything, which is exactly what the Archive plan is for.

Going further

  • Filter the field. /v1/races/search finds races by jurisdiction, run type, going, class, handicap status and field size, with a cursor for large ranges, so a backtest can ask "handicaps on soft ground in Britain" rather than "every race".
  • Skip the maths. The Analyst plan computes strike rate, A/E and level stakes for every horse, trainer and jockey, sliceable by course, distance and going, so the numbers in this guide come back from a single call.
  • Every endpoint, with code: the complete endpoint guide in this repository, and the documentation.

Get a free key · Compare the plans · Read the documentation


The guide and code are MIT licensed. Horse Racing API is run by DeveloperData.net. Questions: https://apihorseracing.com/support.

About

A step-by-step horse racing API tutorial: test whether a betting strategy has ever worked, on real races and the prices that were actually available, using the Horse Racing API as your horse racing data API and nothing but Python

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages