Open-source, self-hosted analytics for AI crawlers, search engines, SEO tools, and automated agents.
See who is crawling your websites, which URLs they request, how often they return, and whether those requests succeed. Track crawler activity across multiple projects from one web dashboard.
Bot Observability provides crawl-side evidence for search and AI visibility work:
- Which AI, search, and SEO crawlers are visiting?
- Which pages are they requesting?
- Are important URLs returning errors or redirects?
- Is crawler activity changing over time?
- Is each project's logging pipeline healthy?
It measures requests reaching your sites, not citations inside AI answers. Use it alongside search and AI visibility tools when you need both crawl activity and citation data.
- Bot Detection — Identifies 130+ bots by User-Agent across AI training crawlers, search engines, SEO tools, social platforms, monitoring services, and CLI tools
- Bot Verification — Confirms bot identity via reverse DNS (PTR) lookups and known IP CIDR ranges
- Multi-Project Support — Track bot traffic across multiple web projects from a single dashboard
- Trend Analysis — Period-over-period comparison (24h / 7d / 30d / 90d / 1y, or a custom date range) showing rising bots, pages, and projects
- Status Quality — Response-code rollups, top failing paths, API/sensitive path hits, and bots with error or UA-only traffic
- AI Crawler Intel — Dedicated view for AI training and search crawlers with confidence breakdowns (verified vs UA-only) and a crawls-vs-visits breakdown by company
- Raw Events — Filterable event log with bot name, path, project, IP, and user-agent details
- Data Health Monitoring — Heartbeat freshness tracking to detect logging pipeline issues
- AI crawler monitoring — Track GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and other AI-related crawlers.
- Technical SEO monitoring — Find crawl errors, redirects, failing paths, and unexpected API or sensitive-route requests.
- GEO research — See which automated agents and retrieval crawlers reach your content as part of broader AI visibility analysis.
- Multi-site reporting — Compare crawler activity across websites, products, environments, or client projects.
- Privacy-conscious observability — Self-host the dashboard and store submitted IPs only as keyed hashes.
The public landing page includes a no-database preview with illustrative data. These captures show the overall product surface and the dashboard’s crawler-mix panel:
- Next.js 16 (App Router, Turbopack)
- TypeScript
- Tailwind CSS v4
- Recharts (chart components)
- PostgreSQL (Aiven or any standard Postgres)
- Node.js 20+
- A PostgreSQL database (Aiven, Neon, or any standard Postgres)
# Install dependencies
npm install
# Set environment variables
cp .env.example .env
# Edit .env with your values| Variable | Required | Description |
|---|---|---|
DATABASE_URL |
Yes | PostgreSQL connection string |
BOT_ADMIN_TOKEN |
Yes | Dashboard login and session-signing secret; use a unique 32+ character value |
BOT_IP_HASH_SECRET |
Yes | Dedicated 32+ character secret for keyed IP hashes; do not reuse another role's secret |
BOT_INGEST_TOKENS |
Yes | JSON object mapping project names to unique 32+ character ingestion keys |
Generate each secret with openssl rand -base64 32 or an equivalent cryptographically secure generator. During migration, the legacy BOT_LOG_TOKEN is accepted only when no new ingestion mapping is configured.
Migrations live in db/migrations/*.sql and are applied in order by scripts/migrate.mjs, which tracks what's already been applied in a schema_migrations table (safe to re-run):
npm run migrate
# or point it at a specific database:
node scripts/migrate.mjs "$DATABASE_URL"The bot_hits_daily and bot_first_seen tables are seeded from existing history by the migration and then kept current by insertHit on every event. Migration 004_weighted_rollups.sql repairs the daily rollup for existing deployments where sampled rows were previously backfilled with unweighted counts. If you apply migrations while an older (pre-rollup) build is still receiving traffic, those rows land in bot_hits but not the rollup; after deploying, run the reconcile script once to rebuild the rollup from raw and restore exact parity (safe to re-run any time you suspect drift):
npm run reconcile-rollups
# or: node scripts/reconcile-rollups.mjs "$DATABASE_URL"npm run devOpen http://localhost:3000. The landing page is at /, the dashboard at /dashboard.
For a first local check, open / before configuring a database. The landing page and dashboard preview use illustrative data; /dashboard requires a configured BOT_ADMIN_TOKEN and DATABASE_URL.
The app is a standard Next.js app and works on Vercel or any Node host that can run next start.
For Vercel:
- Create a PostgreSQL database (Aiven, Neon, or any provider).
- Run
npm run migrate(ornode scripts/migrate.mjs "$DATABASE_URL"). - Generate separate admin, IP-hash, and per-project ingestion secrets.
- Add
DATABASE_URL,BOT_ADMIN_TOKEN,BOT_IP_HASH_SECRET, andBOT_INGEST_TOKENSas environment variables. - Deploy the repository.
- Open
/dashboardand sign in withBOT_ADMIN_TOKEN.
Do not expose any of these secrets in client-side browser code. Tracked sites should send events from a server-only proxy, route, function, or backend logger.
The dashboard has 4 tabs, selected via ?view=:
| Route | Description |
|---|---|
/ |
Landing page with feature overview and login |
/dashboard or /dashboard?view=overview |
Totals, crawler mix, daily trend, time-of-day distribution, movers, and AI crawls-vs-visits by company (default tab) |
/dashboard?view=bots |
Full bot list; add &bot=<name> for a per-bot detail view (trend chart, top pages, first/last seen) or &category=ai to filter to AI bots |
/dashboard?view=health |
2xx/3xx/4xx/5xx mix, failing paths, API/sensitive path hits |
/dashboard?view=events |
Filterable raw event log |
/api/bot-hit |
Authenticated bot event ingestion endpoint |
/login |
POST handler for token auth |
Older URLs (?view=ai, ?view=trends, ?view=status, ?view=pages, ?view=bot&bot=<name>) still work — they 307-redirect to their current equivalent (see src/proxy.ts).
Every view also accepts ?period=, either a preset (1, 7, 30, 90, 365 days) or a custom range as YYYY-MM-DD_YYYY-MM-DD. Periods over 90 days switch views into a rollup-backed "long-range mode" (see Architecture below) that hides path-level panels the daily rollup can't serve.
Tracked sites should POST JSON to the collector's /api/bot-hit endpoint with a project-scoped bearer token or x-bot-log-token header. Bot identity, category, and confidence are derived server-side from the submitted user agent and IP. The collector derives the project from the credential; a new credential cannot submit for another project.
Send status_code when it is available. Status reports show older or incomplete events as not captured; that is not a real HTTP status class.
curl -X POST "$DASHBOARD_URL/api/bot-hit" \
-H "Authorization: Bearer $BOT_INGEST_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"project": "marketing-site",
"environment": "production",
"url": "https://example.com/pricing?ref=ai",
"method": "GET",
"status_code": 200,
"user_agent": "GPTBot/1.0",
"ip": "203.0.113.10",
"referer": "",
"sample_rate": 1
}'Minimal non-blocking Next.js Proxy example:
import { NextResponse, type NextFetchEvent, type NextRequest } from "next/server";
import { isLikelyBotUserAgent } from "@/lib/bots";
function reportBotHit(request: NextRequest) {
return fetch(`${process.env.BOT_OBSERVABILITY_URL}/api/bot-hit`, {
method: "POST",
headers: {
"content-type": "application/json",
authorization: `Bearer ${process.env.BOT_INGEST_TOKEN}`,
},
body: JSON.stringify({
url: request.url,
method: request.method,
status_code: 0,
user_agent: request.headers.get("user-agent") ?? "",
ip: request.headers.get("x-forwarded-for")?.split(",")[0]?.trim() ?? "",
referer: request.headers.get("referer") ?? "",
}),
}).then((response) => {
if (!response.ok) console.error(`[bot-observability] ingestion failed: ${response.status}`);
});
}
export function proxy(request: NextRequest, event: NextFetchEvent) {
const response = NextResponse.next();
const userAgent = request.headers.get("user-agent") ?? "";
if (isLikelyBotUserAgent(userAgent)) {
event.waitUntil(reportBotHit(request).catch(() => undefined));
}
return response;
}Preserve each site's existing redirects, rewrites, locale handling, security headers, and content negotiation around this sender. The collector derives the project from the token, and status_code: 0 means that a pass-through Proxy does not claim to know the final downstream status.
Heartbeat events can be sent periodically to monitor pipeline freshness:
curl -X POST "$DASHBOARD_URL/api/bot-hit" \
-H "Authorization: Bearer $BOT_INGEST_TOKEN" \
-H "Content-Type: application/json" \
-d '{"heartbeat":true,"environment":"production"}'Heartbeats update one row in project_health per project. They are idempotent and do not append rows to bot_hits.
| Field | Required | Notes |
|---|---|---|
project or project_name |
Legacy only | New project-scoped credentials derive this value from the credential; legacy senders default to default |
url |
No | Used to derive host, path, and query_string when those are not provided |
path |
No | Useful when you do not want to send full URLs |
method |
No | Defaults to GET |
status_code or status |
No | Use the final HTTP response status when available |
user_agent |
Yes for bot detection | Non-bot events are ignored unless heartbeat is true |
ip |
No | Enables bot verification; stored only as a keyed HMAC-SHA-256 value, never as the raw submitted IP |
referer |
No | Stored for raw event inspection |
environment |
No | Defaults to production |
is_api_route |
No | Helps the Status tab surface API hits |
sample_rate |
No | Allowed values are 1, 0.5, 0.25, and 0.1; rollups weight each row by the exact integer reciprocal |
heartbeat |
No | Set true for pipeline health events |
The ingestion endpoint enforces a few hardcoded limits (see src/app/api/bot-hit/route.ts):
- Max request body: 32KB (
MAX_BODY_BYTES). Larger requests (byContent-Length) are rejected. - Rate limit: 120 requests/minute per caller IP (
RATE_LIMIT_RPM), tracked in an in-memory, per-serverless-instance store — see the caveat under Architecture; it is not a global ceiling on multi-instance platforms. - Field truncation: string fields are silently truncated, not rejected — 2000 characters for most string fields (
MAX_STRING_LENGTH), 1000 characters forpath(MAX_PATH_LENGTH). Oversized values are cut, not errored. - Responses:
Status Meaning 201Event stored ( { stored: true, bot_name, bot_category, confidence })200Not stored — non-bot, non-heartbeat traffic ( { stored: false, reason: "not_bot" })400Invalid or too-large JSON payload 401Missing/invalid ingestion credential 429Rate limit exceeded 503Ingestion not configured ( DATABASE_URL,BOT_IP_HASH_SECRET, or ingestion credentials missing/weak)
- The dashboard is protected by
BOT_ADMIN_TOKEN, not a full user-management system. A successful login creates a signed, HTTP-only, same-site session cookie valid for 1 year; the token itself is never stored in the cookie. - Ingestion uses project-scoped
BOT_INGEST_TOKENS; keep them server-side and never place them inNEXT_PUBLIC_*variables. - Submitted IP addresses are used for bot verification and then stored only as domain-separated, keyed HMAC-SHA-256 values derived from
BOT_IP_HASH_SECRET. Raw IP storage is not supported. - Rotating
BOT_IP_HASH_SECRETchanges the keyed hash produced for future observations of the same IP. Existing stored hashes remain unchanged. - User agents, paths, referrers, approximate geo fields, deployment URLs, and status codes may be stored.
- Rotate
DATABASE_URLand every role-specific secret before making a previously private deployment public if any may have been exposed outside trusted systems.
- Storage: a single
bot_hitsraw event table, two maintained tables —bot_hits_daily(a(day, project, bot, category, status_class)rollup used for long-range and high-volume views) andbot_first_seen(per-bot first/last-seen timestamps) — plus idempotentproject_healthheartbeat state. Bot events update the raw row and rollups in one transaction; heartbeats update onlyproject_health. If rows are ever ingested by an older build that predates the rollup,npm run reconcile-rollupsrebuilds both tables from raw (see Database Setup). Day buckets are UTC (DATE(created_at)); the UI displays timestamps in Europe/Berlin. - Request-scoped DB client: each request gets its own
postgresclient viacache()+after()(seesrc/lib/db.ts/src/app/dashboard/page.tsx), closed at the end of the request rather than pooled indefinitely — deliberate for small free-tier Postgres connection limits (e.g. Aiven). - Rendering: the dashboard streams server-rendered content with a
Suspenseboundary per view/panel, so slow queries don't block the whole page. - Rate limiting is per-instance, not global:
/api/bot-hit's rate limiter is an in-memoryMapscoped to a single running process (seesrc/app/api/bot-hit/route.ts). On multi-instance serverless platforms like Vercel, each concurrently-running instance enforces its own 120 req/min ceiling independently — there is no shared/global counter. Real aggregate throughput across all instances can therefore be significantly higher than 120 req/min. Do not rely on this limiter as a hard global cap; put a WAF/edge rate limit in front of it if you need one.
npm test/npm run test:unit— unit tests (pure logic: bot detection, category normalization, period parsing, attention-strip thresholds, etc.), no database required.npm run test:integration— integration tests against a real Postgres database, gated onTEST_DATABASE_URLbeing set (skipped otherwise). They apply the migrations, seed fixtures, and clean up after themselves.
Raw events can be retained for a bounded period while daily rollups, first/last-seen data, and project health remain available. The optional cleanup command defaults to 90 days:
npm run retain-raw
# or: node scripts/retain-raw-events.mjs "$DATABASE_URL" 90Schedule it daily or weekly. It only deletes old rows from bot_hits; it does not delete or rebuild rollups, first/last-seen records, or project_health:
DELETE FROM bot_hits
WHERE created_at < now() - interval '90 days';| Category | Description |
|---|---|
| AI Training | Bulk training data collectors (GPTBot, ClaudeBot, etc.) |
| AI Search | Indexers for AI chat products (OAI-SearchBot, PerplexityBot, etc.) |
| AI Agent | On-demand user-triggered fetches (ChatGPT-User, Claude-User, etc.) |
| Search Engine | Traditional search index crawlers (Googlebot, Bingbot, etc.) |
| Social Preview | Link unfurling / share preview cards (Twitterbot, Slack, etc.) |
| SEO Tool | SEO audit and tech detection (Ahrefs, Semrush, etc.) |
| Monitoring | Uptime / performance checks (Pingdom, UptimeRobot, etc.) |
| Archival | Web page preservation (Internet Archive, etc.) |
| Generic / CLI | Uncategorized automated agents (curl, wget, etc.) |

