A lightweight, provider-agnostic multi-agent orchestration framework for Python.
AetherAgents lets you build tool-using agents and compose them into teams, with a
tiny core (only pydantic) and zero required network dependencies. Drop in a
real model via LiteLLM when you're ready — or stay
fully offline with the built-in deterministic MockProvider for tests and demos.
from aetheragents import Agent, MockProvider, tool
@tool()
def add(a: int, b: int) -> int:
"Add two numbers."
return a + b
agent = Agent("calc", MockProvider([[("add", {"a": 2, "b": 2})], "It's 4."]), tools=[add])
print(agent.run("What is 2 + 2?").output) # -> It's 4.| 🪶 Tiny core | The base install depends only on pydantic. Backends (LiteLLM, Anthropic, ChromaDB, FastAPI, OpenTelemetry) are opt-in extras. |
| 🔌 Provider-agnostic | One LLMProvider interface. LiteLLMProvider reaches 100+ models, AnthropicProvider talks to Claude natively, MockProvider runs offline. |
| 🌊 Streaming | agent.astream() yields live text deltas and tool events; every provider streams (native or fallback). |
| 📦 Structured output | run(prompt, response_model=MyModel) returns a validated Pydantic instance, with automatic schema-violation retries. |
| 🛠️ Real tool schemas | The @tool decorator generates JSON-Schema from your function signature and type hints — no hand-written specs. Built-in calculator / clock / HTTP tools included. |
| 🤝 Multi-agent | Sequential, parallel, route, debate, map-reduce, and agent-as-tool delegation. |
| 🧠 Memory & sessions | Rolling window + long-term vector recall, JSON-persisted Sessions, and session forks for branching. |
| 🛡️ Guardrails & budgets | Input/output hooks plus hard cost budgets (max_cost_usd) and human tool approval. |
| 💰 Cost tracking | result.cost_usd estimates spend per run; export full traces as JSON. |
| ✅ Testable | Deterministic mock provider, offline eval harness, and a full unit suite — no API keys. |
| 🔭 Observable & resilient | Optional OpenTelemetry, step traces, RetryingProvider backoff, Mermaid team diagrams. |
| ⌨️ CLI | aetheragents run, version, and doctor for offline demos and install checks. |
pip install aetheragents # tiny core (pydantic only)
pip install 'aetheragents[litellm]' # 100+ models via LiteLLM
pip install 'aetheragents[anthropic]' # Claude via the official Anthropic SDK
pip install 'aetheragents[chroma]' # ChromaDB-backed long-term memory
pip install 'aetheragents[server]' # FastAPI HTTP server
pip install 'aetheragents[all,dev]' # everything + test/lint toolingUntil it's published to PyPI, install from source:
pip install -e '.[dev]'.
An Agent runs a reasoning loop: call the model, execute any requested tools,
feed results back, repeat until the model answers or max_steps is hit. Every
run returns a structured AgentResult with the final output, a step-by-step
trace and token usage.
from aetheragents import Agent, LiteLLMProvider, tool
@tool()
def get_weather(city: str) -> str:
"Get the current weather for a city."
return f"It's sunny in {city}."
agent = Agent(
"assistant",
LiteLLMProvider("gpt-4o"), # needs: pip install 'aetheragents[litellm]'
instructions="You are a helpful assistant.",
tools=[get_weather],
max_steps=6,
)
result = agent.run("What's the weather in Melbourne?")
print(result.output)
for step in result.steps:
print(step.type, step.name or "", step.content or step.arguments)Prefer Claude natively? Swap the provider:
from aetheragents import AnthropicProvider # needs: pip install 'aetheragents[anthropic]'
agent = Agent("assistant", AnthropicProvider("claude-sonnet-5"), tools=[get_weather])Wrap any provider for resilience:
from aetheragents import RetryingProvider
provider = RetryingProvider(AnthropicProvider(), max_retries=3) # exponential backoffastream() exposes the whole run as live events — text deltas, tool calls,
tool results, and a terminal result:
async for event in agent.astream("What's the weather in Melbourne?"):
if event.type == "delta":
print(event.delta, end="", flush=True) # tokens as they arrive
elif event.type == "step" and event.step.type == "tool_call":
print(f"\n[calling {event.step.name}...]")
elif event.type == "result":
final = event.result # full AgentResultAsk for a Pydantic model and get a validated instance back. Invalid answers are automatically fed back to the model for correction:
from pydantic import BaseModel
class CityFacts(BaseModel):
city: str
country: str
population_millions: float
result = agent.run("Tell me about Melbourne", response_model=CityFacts)
result.parsed.country # -> "Australia", guaranteed schema-validDecorate any function with @tool(). The argument schema is derived from the
signature — required vs optional, and JSON types from your annotations.
@tool(description="Search the knowledge base.")
def search(query: str, limit: int = 5) -> list[str]:
...
# -> parameters: {query: string (required), limit: integer}Or start from the built-ins — a safe AST-based calculator (no eval), a UTC
clock, and a capped HTTP fetcher:
from aetheragents import builtin_tools
agent = Agent("helper", provider, tools=builtin_tools())Need file access? file_tools confines reads/writes to a sandbox directory
(traversal, absolute paths and symlink escapes are rejected):
from aetheragents import file_tools
agent = Agent("writer", provider, tools=file_tools("./workspace"))
agent = Agent("reader", provider, tools=file_tools("./docs", readonly=True))Validate or transform what goes into and comes out of an agent. A guardrail is
any (str) -> str callable; raise GuardrailError to block the run:
from aetheragents import Agent, blocklist, max_length, redact
agent = Agent(
"support",
provider,
input_guardrails=[max_length(2000)], # reject huge prompts
output_guardrails=[redact(["hunter2"]), blocklist(["ssn"])],
)Every result reports the model used and an estimated cost (or None for
unknown models). Cap spend per agent or per run — exceeding the budget raises
BudgetExceeded (unknown models count as $0 so MockProvider stays offline-friendly):
from aetheragents import Agent, BudgetExceeded, register_model_cost
agent = Agent("assistant", provider, max_cost_usd=0.05)
try:
result = agent.run("summarise this")
print(result.model, result.cost_usd)
except BudgetExceeded as e:
print(f"stopped at ${e.spent:.4f} / ${e.budget}")
register_model_cost("my-local-model", 0.0, 0.0) # $/MTok input, output
result.export_trace("traces/last-run.json") # full step + message dumpGate tool execution with a callable (CLI prompt, policy engine, …). Default
None auto-approves everything. When the model returns multiple tool calls in
one step, set parallel_tools=True to run them concurrently:
def approve(name: str, args: dict) -> bool:
return name != "delete_everything"
agent = Agent(
"safe",
provider,
tools=[...],
tool_approval=approve, # False -> "User denied tool execution"
parallel_tools=True, # asyncio.gather independent tool calls
)import asyncio
from aetheragents import Agent, MockProvider, Orchestrator, keyword_router
team = Orchestrator([
Agent("researcher", MockProvider(handler=lambda m: "facts...")),
Agent("writer", MockProvider(handler=lambda m: "# Article")),
Agent("critic", MockProvider(handler=lambda m: "push back...")),
Agent("judge", MockProvider(handler=lambda m: "final synthesis")),
])
# Pipeline: researcher's output feeds the writer
final = asyncio.run(team.sequential("Write about X", order=["researcher", "writer"]))[-1]
# Fan-out: run everyone on the same task
results = asyncio.run(team.parallel("Summarise X"))
# Route: pick one agent by keyword
router = keyword_router({"write": "writer"}, default="researcher")
chosen = asyncio.run(team.route("please write a post", selector=router))
# Debate: multi-round discussion + optional synthesizer
debate = asyncio.run(team.debate(
"Should we ship Friday?",
agents=["writer", "critic"],
rounds=2,
synthesizer="judge",
))
print(debate["output"])
# Map-reduce: parallel workers, then a reducer over their outputs
mr = asyncio.run(team.map_reduce(
"Research topic X",
worker_names=["researcher", "critic"],
reducer_name="writer",
))
print(mr["output"])
# Docs: Mermaid diagram of the team
print(team.to_mermaid())Delegation — expose any agent as a tool so a "manager" agent can call it:
manager = Agent("manager", provider, tools=[team["researcher"].as_tool()])from aetheragents import Session
session = Session("support-42", path="sessions/support-42.json")
agent.run("My printer is on fire", session=session)
# Branch the conversation without mutating the parent history
branch = session.fork(name="try-reset")
agent.run("Have you tried turning it off and on?", session=branch)from aetheragents import Agent, MockProvider
from aetheragents.eval import run_cases
agent = Agent("demo", MockProvider(["hello world", "42"]))
report = run_cases(agent, [
{"prompt": "greet", "expect_contains": "hello"},
{"prompt": "answer", "expect_contains": "42"},
])
assert report.okaetheragents version
aetheragents doctor # which extras are installed?
aetheragents run --agent demo "hello" # offline MockProvider demofrom aetheragents import Agent, MemoryManager, MockProvider
mem = MemoryManager("assistant") # in-memory vector store by default
agent = Agent("assistant", MockProvider(["ok"]), memory=mem)
agent.run("Remember: my favourite colour is teal.")
# Later runs automatically recall relevant past messages.Use ChromaDB for persistence:
from aetheragents import MemoryManager
from aetheragents.core.memory import ChromaVectorStore # needs [chroma]
mem = MemoryManager("assistant", vector_store=ChromaVectorStore("assistant"))from aetheragents import Agent, MockProvider
from aetheragents.server import create_app # needs [server]
app = create_app({"echo": Agent("echo", MockProvider(default="hi"))})
# uvicorn mymodule:app
# POST /agents/echo/run {"prompt": "..."} -> JSON result (+ model, cost_usd)
# POST /agents/echo/stream {"prompt": "..."} -> Server-Sent Events (delta/step/result) ┌─────────────────────────────────────────────┐
│ Orchestrator │
│ sequential · parallel · route · debate · map-reduce │
└───────────────┬──────────────┬──────────────┘
│ │
┌───────▼──────┐ ┌─────▼────────┐
│ Agent │ │ Agent │ reasoning loop
└───┬──────┬───┘ └──────────────┘
│ │
┌─────────────▼─┐ ┌─▼──────────────┐ ┌───────────────────────────┐
│ ToolRegistry │ │ MemoryManager │ │ LLMProvider │
│ (auto schema) │ │ short + vector │ │ Mock | LiteLLM | Anthropic│
│ + built-ins │ │ + Session │ │ (+ RetryingProvider wrap) │
└───────────────┘ └────────────────┘ └───────────────────────────┘
pip install -e '.[dev]'
pytest # fully offline
ruff check . # lint
aetheragents doctorRun the examples:
python examples/quickstart.py
python examples/research_team.py
python examples/streaming_and_structured.py- Pluggable embedding backends for
MemoryManager - OpenTelemetry span coverage for tools and providers
- Richer CLI (session resume, team run)
See CHANGELOG.md for release history.
MIT © Sebby1770