Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

16 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

four

Four functions compose. The loop is the evaluator.

An agent that runs bash commands. A generator that produces notebooks. A tool that writes Python code. Same loop. Different functions.

from four import run, litellm_invoke, regex_parse, local_env, save_trajectory

run(
    G=litellm_invoke("anthropic/claude-sonnet-4-5-20250929"),
    V1=regex_parse(),
    V2=local_env(),
    emit=save_trajectory(),
    system="You are a helpful assistant that executes bash commands.",
    prompt="Find all Python files in /tmp and count lines in each",
    max_steps=50,
)

The algebra

invoke   : G   -- messages → Result[raw]
parse    : V1  -- raw → Result[list[action]]
validate : V2  -- action → Result[observation | Exit]
emit     : IO  -- (messages, outcome) → Path

The loop chains them: (G → V1 → [V2, V2, ...])* → emit

Each step: G queries the model, V1 extracts all actions, V2 executes each one. If V1 fails, the error becomes a user message and the loop continues — the model sees its mistake and self-corrects on the next turn. Four functions compose.

The loop

The entire evaluator is 22 lines:

def run(G, V1, V2, emit, system, prompt, max_steps=100, max_format_errors=3):
    messages = [{"role": "system", "content": system},
                {"role": "user", "content": prompt}]
    consecutive_format_errors = 0

    for step in range(max_steps):
        raw = G(messages)
        if isinstance(raw, Err):
            return emit(messages, f"model_error: {raw.error}")

        actions = V1(raw.value)
        if isinstance(actions, Err):
            consecutive_format_errors += 1
            if 0 < max_format_errors <= consecutive_format_errors:
                return emit(messages, f"repeated_format_error: {actions.error}")
            messages.append({"role": "user", "content": f"Format error: {actions.error}..."})
            continue

        consecutive_format_errors = 0
        for action in actions.value:
            result = V2(action)
            if isinstance(result, Err):
                return emit(messages, result.error)
            messages.append(result.value)

    return emit(messages, "max_steps_reached")

That's it. No config files. No YAML. No SDK. No Pydantic models. No Jinja2 templates baked into the code. Just four functions that take and return well-typed values, chained in a loop.

What it replaces

mini-swe-agent four
YAML config with 40+ parameters Four function arguments
Pydantic model configs Plain functions
Jinja2 templates in config Templates passed as strings
FormatError + InterruptAgentFlow hierarchy Ok | Err
Inner retry loop for format errors Format error as user message, outer loop continues
1000+ lines of boilerplate 22-line loop

Same capability. Different shape.

Components

G — invoke. Queries the LLM. Returns Ok(text) or Err(reason).

  • litellm_invoke() — plain text with markdown code blocks
  • litellm_toolcall_invoke() — structured tool calls
  • retry_invoke(fn) — wraps any G with exponential backoff retry

V1 — parse. Extracts actions from raw output. Returns Ok(list[command]) or Err(reason).

  • regex_parse() — extracts mswea_bash_command, bash, or ```sh blocks (returns all matches)
  • toolcall_parse() — parses JSON tool call payloads (returns all bash commands)

V2 — validate. Executes each action. Returns Ok(observation) or Err(exit).

  • local_env() — subprocess execution with output truncation and exit signal detection

emit — IO. Saves the trajectory. Returns Path.

  • save_trajectory() — JSON files with outcome and full message history

Extending

Every component is swappable. The loop doesn't care:

# Tool-calling instead of regex
G=litellm_toolcall_invoke("openai/gpt-4o"),
V1=toolcall_parse(),

# Container execution instead of local
V2=docker_env(image="python:3.12"),

# Retry on transient errors
G=retry_invoke(litellm_invoke("openai/gpt-4o")),

# Abort after 2 format errors instead of 3
max_format_errors=2,

The same algebra, different domains

The four-function loop generates notebooks, agents, Python code, and CLI tools. The only difference is what V2 validates and what emit produces:

Domain V2 validates emit produces
Notebooks AST + chart execution .ipynb with embedded PNGs
Agents bash execution JSON trajectory
Python code type checking .py files
CLI tools compilation binary + man page

The loop doesn't know what it's evaluating. It only chains Result types.

Why four?

Four is the minimum. Remove any one and the loop breaks:

  • No G → nothing to evaluate
  • No V1 → can't extract actions from raw text
  • No V2 → can't execute or observe
  • No emit → can't persist results

Format recovery is built into the loop — no separate function needed.

Philosophy

The framework doesn't call itself category theory. It calls itself algebra. Four functions compose. The loop is the evaluator.

#agenticcoding #functional-programming #python #llm #agents #monads

About

functions compose. loop is the evaluator.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages