This document covers starting, managing, and troubleshooting the Agentlip hub daemon.
- Starting the Hub
- Startup Sequence
- Recovery and Restart
- Migrations
- Doctor Command
- Configuration
- Graceful Shutdown
agentlipd up [--workspace <path>] [--host 127.0.0.1] [--port 0] [--idle-shutdown-ms <ms>] [--json]Options:
--workspace: Workspace root directory (default: auto-discover fromcwd; initializes atcwdif none found)--host: Bind address (default:127.0.0.1). Localhost-only is enforced. See Security: Transport Security.--port: Port number (default:0= random available port)--idle-shutdown-ms: Optional idle auto-shutdown timeout (milliseconds, daemon mode only — requiresworkspaceRoot/--workspace). When enabled, the hub will stop if there are no WS clients and no HTTP activity for the configured duration. Note:GET /healthdoes not reset the idle timer.--json: Output connection info as JSON (never prints auth token)
Output (human):
✓ Hub started
Host: 127.0.0.1
Port: 54321
Workspace: /path/to/workspace
Instance: 8a7f3e2d-...
Output (--json):
{
"status": "running",
"host": "127.0.0.1",
"port": 54321,
"workspace_root": "/path/to/workspace",
"instance_id": "8a7f3e2d-..."
}Exit codes:
0: Clean shutdown1: Error10: Writer lock conflict (hub already running)
Auth token & server.json:
- In daemon mode, the hub persists connection info (including the auth token) to
.agentlip/server.jsonwith mode 0600. - The auth token is never printed to stdout/stderr and must not appear in process argv.
From TypeScript/JavaScript:
import { startHub } from "@agentlip/hub";
const hub = await startHub({
host: "127.0.0.1",
port: 0, // random port
workspaceRoot: process.cwd(),
enableFts: true, // optional full-text search
});
console.log(`Hub running on ${hub.host}:${hub.port}`);
// Later: graceful shutdown
await hub.stop();Reference: packages/hub/src/index.ts lines 150-470
The hub follows a deterministic startup sequence to ensure safe initialization:
File: packages/workspace/src/index.ts lines 40-100
- Auto-discovers workspace by walking upward from current directory
- Stops at filesystem boundary or user home directory (security boundary)
- Looks for
.agentlip/db.sqlite3marker file
File: packages/kernel/src/index.ts lines 25-70
const db = openDb({ dbPath: ".agentlip/db.sqlite3" });PRAGMAs applied:
journal_mode = WAL(Write-Ahead Logging for crash recovery)foreign_keys = ON(referential integrity)busy_timeout = 5000(5s wait for lock contention)synchronous = NORMAL(balance safety/performance)
File: packages/kernel/src/index.ts lines 130-230
runMigrations({
db,
migrationsDir: "./migrations",
enableFts: true, // optional
});Migration files:
0001_schema_v1.sql- Core schema (channels, topics, messages, events, attachments, enrichments)0001_schema_v1_fts.sql- Optional FTS5 full-text search index (opportunistic, non-fatal)
Tracking:
- Current schema version stored in
meta.schema_version - Migrations are forward-only (no rollbacks)
- Before migration: creates timestamped backup (
.backup-v{version}-{timestamp})
File: packages/kernel/src/index.ts lines 75-115
Ensures meta table has required keys:
db_id: UUIDv4 (generated once at init, never changes)schema_version: Current version (initially '1')created_at: ISO8601 timestamp
File: packages/hub/src/lock.ts lines 25-90
Daemon mode only (when workspaceRoot is provided):
.agentlip/locks/writer.lock
Lock contents:
{pid}\n{started_at}
Stale lock detection:
- Read
server.jsonto get hub instance info - Check
/healthendpoint to verify hub is responsive - If health check fails → lock is stale, remove and retry acquisition
- Max retries: 3 (with 100ms delay between attempts)
Reference: packages/hub/src/lock.ts lines 45-120
File: packages/hub/src/authToken.ts lines 10-20
const authToken = generateAuthToken(); // 32 bytes (256 bits) cryptographically randomStored in: .agentlip/server.json (mode 0600, owner read/write only)
Token never logged in error messages or structured logs.
File: packages/hub/src/index.ts lines 250-350
const server = Bun.serve({
hostname: host,
port: port,
fetch: handleRequest,
websocket: wsHandlers,
});Endpoints:
GET /health- Unauthenticated health checkGET /ws- WebSocket upgrade (requires auth token in query param:?token=...)/api/v1/*- HTTP API (mutations requireAuthorization: Bearer <token>)
File: packages/hub/src/serverJson.ts lines 40-85
Daemon mode only:
{
"instance_id": "8a7f3e2d-...",
"db_id": "f3e2d4b1-...",
"port": 54321,
"host": "127.0.0.1",
"auth_token": "a1b2c3d4...",
"pid": 12345,
"started_at": "2026-02-05T18:30:00.000Z",
"protocol_version": "1",
"schema_version": 1
}Atomic write:
- Write to temp file:
.server.json.tmp.{random} - Set mode 0600
- Rename to
server.json(atomic on same filesystem) - Verify final permissions
Security: Mode 0600 ensures only workspace owner can read auth token.
File: packages/hub/src/config.ts lines 20-120
Daemon mode only:
const config = await loadWorkspaceConfig(workspaceRoot);Config file: agentlip.config.ts in workspace root
Contents:
export default {
plugins: [
{
type: "linkifier",
name: "url-preview",
module: "./plugins/url-preview.ts",
enabled: true,
config: { /* plugin-specific */ },
},
],
};Note: Config load failure aborts startup (releases lock before exiting).
File: packages/hub/src/index.ts lines 380-440
If workspace config declares enabled plugins:
- Linkifier plugins registered for
message.created/message.editedevents - Extractor plugins registered for same events
- Execution is asynchronous (doesn't block HTTP response)
- Failures logged but don't affect message ingestion
Agentlip uses SQLite WAL mode for automatic crash recovery:
How WAL works:
- Writes go to
.db.sqlite3-wal(Write-Ahead Log) file - On clean shutdown: WAL checkpointed back to main DB file
- On crash: next startup replays WAL automatically
Recovery steps:
- Open database (SQLite replays WAL if present)
- Database is now in consistent state as of last committed transaction
- In-flight transactions during crash are rolled back
WAL checkpoint:
On graceful shutdown (packages/hub/src/index.ts lines 520-530):
db.run("PRAGMA wal_checkpoint(TRUNCATE)");TRUNCATE mode:
- Checkpoints all WAL frames back to main DB
- Truncates WAL file to reclaim disk space
- Best-effort (failure logged but doesn't block shutdown)
File: packages/hub/src/lock.ts lines 100-145
Scenario: Hub crashed without cleaning up writer.lock
Detection:
- Next startup attempts lock acquisition
- Finds existing lock file
- Reads
server.jsonto get previous instance info - Calls health check:
GET http://{host}:{port}/health - Verifies
instance_idmatches (ensures same hub instance, not new hub on recycled port)
If health check fails:
- Lock is stale → remove lock file
- Retry acquisition (max 3 attempts)
If health check succeeds:
- Lock is live → abort startup with error:
Writer lock already held by live hub. Cannot start another hub instance.
File: packages/hub/src/serverJson.ts lines 90-120
On startup:
- Check if
server.jsonexists - If exists: attempt health check (same as lock staleness detection)
- If stale: overwrite with new instance info
- If live: abort startup
On clean shutdown:
- Remove
server.json:packages/hub/src/serverJson.tslines 125-140 - Remove
writer.lock:packages/hub/src/lock.tslines 150-165
Directory: migrations/ (repo root)
Files:
0001_schema_v1.sql- Core schema (required)0001_schema_v1_fts.sql- Full-text search index (optional)
Automatic: Migrations run on every hub startup (packages/kernel/src/index.ts lines 160-230)
Idempotent: Migration SQL uses CREATE TABLE IF NOT EXISTS and CREATE INDEX IF NOT EXISTS
Tracking:
SELECT value FROM meta WHERE key = 'schema_version';
-- Returns: '1'Optional feature controlled by:
- Explicit
enableFtsoption instartHub() - Environment variable:
AGENTLIP_ENABLE_FTS=1
Resolution order (packages/hub/src/index.ts lines 200-215):
function resolveFtsEnabled(enableFts?: boolean): boolean {
if (enableFts !== undefined) return enableFts; // 1. Explicit option
const env = process.env.AGENTLIP_ENABLE_FTS;
if (env === "1") return true; // 2. Env var
if (env === "0") return false;
return false; // 3. Default: disabled
}Non-fatal:
- If FTS migration fails (e.g., SQLite build without FTS5 support), hub continues startup
- Error logged but doesn't abort initialization
- Search commands will return error if FTS unavailable
File: packages/kernel/src/index.ts lines 115-135
Automatic backup:
.agentlip/db.sqlite3.backup-v{fromVersion}-{timestamp}
Example:
.agentlip/db.sqlite3.backup-v0-2026-02-05T15-30-45-123Z
.agentlip/db.sqlite3-wal.backup-v0-2026-02-05T15-30-45-123Z
Restoration:
# Stop hub first
mv .agentlip/db.sqlite3.backup-v0-2026-02-05T15-30-45-123Z .agentlip/db.sqlite3
mv .agentlip/db.sqlite3-wal.backup-v0-2026-02-05T15-30-45-123Z .agentlip/db.sqlite3-wal
# Restart hub
agentlipd upCLI: agentlip doctor [--workspace /path] [--json]
File: packages/cli/src/agentlip.ts lines 150-220
- Workspace discovery (upward walk from current directory)
- Database file exists (
.agentlip/db.sqlite3) - Database can be opened (read-only mode)
- Schema version (from
meta.schema_version) - Query-only mode (whether DB is read-only)
- Database ID (from
meta.db_id)
Human-readable:
✓ Workspace found
Workspace Root: /Users/alice/my-project
Database Path: /Users/alice/my-project/.agentlip/db.sqlite3
Database ID: f3e2d4b1-c9a7-8e5f-6d2b-1a4c3e7f9b2d
Schema Version: 1
Query Only: no
JSON mode:
{
"status": "ok",
"workspace_root": "/Users/alice/my-project",
"db_path": "/Users/alice/my-project/.agentlip/db.sqlite3",
"db_id": "f3e2d4b1-c9a7-8e5f-6d2b-1a4c3e7f9b2d",
"schema_version": 1,
"query_only": false
}Error (workspace not found):
{
"status": "error",
"error": "No workspace found (no .agentlip/db.sqlite3 in directory tree starting from /Users/alice/random-dir)"
}0- OK1- Error (workspace not found, DB issues, etc.)
File: agentlip.config.ts in workspace root
Loaded by: packages/hub/src/config.ts lines 20-120
Schema:
export interface WorkspaceConfig {
plugins: PluginConfig[];
}
export interface PluginConfig {
type: "linkifier" | "extractor";
name: string;
module: string; // relative path from workspace root
enabled: boolean;
config?: Record<string, unknown>; // plugin-specific config
timeout_ms?: number; // default: 5000
circuit_breaker?: {
failure_threshold?: number; // default: 3
cooldown_ms?: number; // default: 60000
};
}Example:
export default {
plugins: [
{
type: "linkifier",
name: "url-preview",
module: "./plugins/url-preview.ts",
enabled: true,
config: {
max_urls: 5,
timeout_ms: 3000,
},
},
{
type: "extractor",
name: "jira-tickets",
module: "./plugins/jira-extractor.ts",
enabled: false, // disabled; won't execute
},
],
};Dynamic reload:
- v1: Config loaded once at startup
- Future: Consider SIGHUP or file watcher for hot reload
File: packages/hub/src/rateLimiter.ts lines 1-200
Defaults (packages/hub/src/rateLimiter.ts lines 125-130):
{
perClient: { limit: 100, windowMs: 1000 }, // 100 req/s per client
global: { limit: 1000, windowMs: 1000 }, // 1000 req/s global
}Configure via startHub options:
await startHub({
rateLimitPerClient: { limit: 50, windowMs: 1000 },
rateLimitGlobal: { limit: 500, windowMs: 1000 },
disableRateLimiting: false, // disable for testing
});Headers in responses:
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 95
X-RateLimit-Reset: 1704477890
429 Too Many Requests:
{
"error": "Rate limit exceeded",
"code": "RATE_LIMITED",
"details": {
"limit": 100,
"window": "1s",
"retry_after": 5
}
}Cleanup:
- Expired rate limit buckets cleaned every 60s
- Automatic cleanup stops on graceful shutdown
File: packages/hub/src/index.ts lines 500-570
await hub.stop();Steps:
-
Set shutdown flag (reject new non-health requests)
- New requests get 503 response:
{ code: "SHUTTING_DOWN" } - Health endpoint continues responding
- New requests get 503 response:
-
Wait for in-flight requests (max 10s)
- Track pending requests via
inflightPromisesset - Drain with timeout:
Promise.race([drain, timeout(10000)])
- Track pending requests via
-
Close WebSocket connections (code 1001 = going away)
- Hub broadcasts close to all connected clients
- Clients should reconnect after delay
-
Stop rate limiter cleanup
- Clear cleanup interval
-
Attempt server stop (bounded wait)
- Call
server.stop(true)(force close) - Known issue (Bun 1.3.x):
stop()can hang after WebSocket connections - Mitigation: Race stop against 250ms timeout; proceed with cleanup anyway
- Call
-
WAL checkpoint (reclaim disk space)
db.run("PRAGMA wal_checkpoint(TRUNCATE)");
- Best-effort; failure logged but doesn't block shutdown
-
Close database
db.close();
-
Remove server.json (daemon mode only)
await removeServerJson({ workspaceRoot });
-
Release writer lock (daemon mode only)
await releaseWriterLock({ workspaceRoot });
CLI daemon:
SIGINT(Ctrl+C) → graceful shutdownSIGTERM→ graceful shutdown
Unhandled signals:
SIGKILL→ immediate termination (no cleanup; lock/server.json left behind; next startup detects stale files)
Total shutdown time bounded:
- In-flight drain: 10s
- Server stop: 250ms
- Total: ~10.5s worst-case
After timeout:
- Proceeds with cleanup regardless of drain status
- Outstanding requests may be aborted
File: packages/hub/src/index.ts lines 25-90
Format: JSON lines to stdout
Example:
{
"ts": "2026-02-05T18:30:45.123Z",
"level": "info",
"msg": "request",
"method": "POST",
"path": "/api/v1/messages",
"status": 201,
"duration_ms": 15,
"instance_id": "8a7f3e2d-...",
"request_id": "a1b2c3d4-...",
"event_ids": [42, 43]
}Suppressed in test environments:
NODE_ENV=testVITESTorJEST_WORKER_IDset- Entry point matches
*.test.{js,ts}
Never logged:
- Auth tokens (full or partial)
- Full message content (only message IDs)
- User credentials
Endpoint: GET /health
Unauthenticated (always accessible, even during shutdown)
Response:
{
"status": "ok",
"instance_id": "8a7f3e2d-4b1c-9d6e-5f8a-7c9e2b3d4a5f",
"db_id": "f3e2d4b1-c9a7-8e5f-6d2b-1a4c3e7f9b2d",
"schema_version": 1,
"protocol_version": "1",
"pid": 12345,
"uptime_seconds": 3600
}Use cases:
- Stale lock detection during startup
- Load balancer health checks (if exposed via reverse proxy)
- Monitoring systems
Header: X-Request-ID
Behavior:
- Client can provide
X-Request-IDin request header - If not provided, hub generates random UUIDv4
- Returned in response header
- Included in structured logs (for request tracing)
Error:
Writer lock already held by live hub. Cannot start another hub instance.
Diagnosis:
# Check if hub is actually running
ps aux | grep agentlipd
# Check lock file
cat .agentlip/locks/writer.lock
# Output: {pid}\n{timestamp}
# Verify hub health
curl http://127.0.0.1:{port}/healthResolution:
- If hub is running: stop it first (
kill {pid}or Ctrl+C) - If hub crashed: remove stale lock manually
rm .agentlip/locks/writer.lock
- Restart hub:
agentlipd up
Error:
database is locked
Causes:
- Another process has exclusive lock (unlikely with WAL mode)
- Long-running transaction holding write lock
- Filesystem issues (NFS, network drive)
Diagnosis:
# Check for other processes with DB open
lsof .agentlip/db.sqlite3
# Check WAL file size (large WAL indicates checkpoint issues)
ls -lh .agentlip/db.sqlite3-walResolution:
- Stop all hub instances
- Checkpoint WAL manually:
sqlite3 .agentlip/db.sqlite3 "PRAGMA wal_checkpoint(TRUNCATE);" - Restart hub
Error:
Migration file not found: /path/to/migrations/0001_schema_v1.sql
Resolution:
- Ensure
migrationsDirpoints to correct path - Default:
{repo}/migrations/ - Programmatic start: provide explicit
migrationsDiroption
Error:
Migration failed: table already exists
Resolution:
- Migration SQL should use
IF NOT EXISTSfor idempotency - Manual fix: inspect schema version, manually apply missing parts
- Last resort: restore from backup, rerun migrations
Error (in CLI):
Full-text search not available: messages_fts table does not exist.
Enable FTS by running migrations with enableFts=true.
Resolution:
- Stop hub
- Enable FTS:
AGENTLIP_ENABLE_FTS=1 agentlipd up
- Or programmatic:
await startHub({ enableFts: true });
Check FTS status:
import { isFtsAvailable } from "@agentlip/kernel";
const hasFts = isFtsAvailable(db);Check circuit breaker state:
Circuit breaker opens after 3 failures (default). Plugin skipped for 60s cooldown.
Logs:
{
"ts": "2026-02-05T18:30:45.123Z",
"level": "warn",
"msg": "[plugins] linkifier pipeline failed for message msg-123",
"error": "Plugin timed out after 5000ms"
}Resolution:
- Check plugin code for infinite loops or blocking I/O
- Increase timeout in config:
{ timeout_ms: 10000, // 10s }
- Fix plugin code and restart hub (circuit auto-resets after cooldown)
WAL checkpoint frequency:
# Manual checkpoint
sqlite3 .agentlip/db.sqlite3 "PRAGMA wal_checkpoint(TRUNCATE);"
# Autocheckpoint (default: every 1000 pages)
sqlite3 .agentlip/db.sqlite3 "PRAGMA wal_autocheckpoint=1000;"Analyze query performance:
sqlite3 .agentlip/db.sqlite3 "ANALYZE;"WebSocket connections:
- Default: no explicit limit (OS default: ~1024 on Linux)
- Future: add
maxConnectionsoption
HTTP requests:
- Rate limited (see Configuration: Rate Limiting)
Timeout adjustment:
- Default: 5s
- Increase for slow plugins (API calls, etc.)
- Monitor timeout errors in logs
Parallelism:
- Plugins execute sequentially per message
- Multiple messages processed concurrently
- Consider async batching for bulk enrichment (future feature)
# Stop hub first (ensures clean state)
agentlipd down
# Backup database + WAL
cp .agentlip/db.sqlite3 backups/db-$(date +%Y%m%d-%H%M%S).sqlite3
cp .agentlip/db.sqlite3-wal backups/db-$(date +%Y%m%d-%H%M%S).sqlite3-wal
# Restart hub
agentlipd up# While hub is running
sqlite3 .agentlip/db.sqlite3 ".backup backups/db-$(date +%Y%m%d-%H%M%S).sqlite3"WAL mode allows hot backup (consistent snapshot without stopping hub).
# Stop hub
agentlipd down
# Restore database
cp backups/db-20260205-153045.sqlite3 .agentlip/db.sqlite3
# Restart hub (will replay WAL if present)
agentlipd upThe hub UI was migrated from inline HTML/CSS/JS to a Svelte 5 SPA served from /ui/* routes. The migration used a temporary feature flag (HUB_UI_SPA_ENABLED) during rollout; that flag was removed at final cutover.
- Gate 1-3 (completed): SPA build pipeline, bootstrap endpoint, route migration with feature flag
- Soak period (completed): Flag defaulted to
true; legacy code path available viaHUB_UI_SPA_ENABLED=false - Gate 4 (completed): Legacy code removed; CSP tightened; flag removed
After legacy code removal (current state):
- Rollback method: Redeploy previous known-good release (e.g., v0.X.Y-1)
- Feature flag no longer available: Setting
HUB_UI_SPA_ENABLEDhas no effect - CSP updated: Removed
'unsafe-inline'fromscript-srcandstyle-src(SPA has no inline scripts/styles)
Cutover verification checklist (required before/at removal):
- Run
bun run typecheck - Run
bun test packages/hub - Run full workspace
bun test - Verify route matrix behavior:
/uiand deep/ui/*routes serve SPA shell/ui/bootstrapreturns runtime JSON/ui/assets/*serves assets and missing assets return 404 (no shell fallback)- no-auth mode returns 503 for all
/ui/*
- Verify CSP headers include no
unsafe-inlineinscript-src/style-src - Document rollback path for production deployments (redeploy previous release)
Steady-state behavior:
| Route | Behavior |
|---|---|
/ui |
Serves SPA shell (Svelte app) |
/ui/bootstrap |
Returns runtime config JSON (baseUrl, wsUrl, authToken) |
/ui/assets/* |
Serves static JS/CSS assets (hashed files get immutable cache) |
Deep client routes (/ui/channels/:id, /ui/topics/:id, /ui/events) |
SPA shell (client-side routing) |
Missing /ui/assets/* |
404 (never SPA fallback) |
All /ui/* routes (no auth token) |
503 (UI unavailable) |
Emergency rollback procedure:
If critical UI regression discovered post-cutover:
- Identify last known-good release (e.g., v0.1.5)
- Deploy previous release:
git checkout v0.1.5 bun install agentlipd down agentlipd up
- Or via package manager (if using published releases):
bun add @agentlip/hub@0.1.5 agentlipd down agentlipd up
- Report issue and prepare patch release
CSP policy (current):
default-src 'self';
script-src 'self';
style-src 'self';
connect-src 'self' ws://localhost:* ws://127.0.0.1:*;
frame-ancestors 'none'
Note: 'unsafe-inline' removed from script-src and style-src after legacy removal (SPA uses external JS/CSS files only).
- Security Documentation - Threat model, authentication, plugin isolation
- API Reference - HTTP and WebSocket API endpoints
- Plugin Development - Writing custom plugins
- AGENTLIP_PLAN.md - Full system specification