This system lets an AI agent execute user requests across whatever booking domains and transactional tools it has available, end to end. The agent interprets intent, discovers options, produces a proposed plan or sequence of composable actions, obtains explicit user confirmation for any state-changing action, executes against pluggable provider backends, and records every step in an append-only log.
The design assumes the agent will make mistakes, providers will fail, and users will change their minds. Its job is to make those outcomes recoverable rather than to pretend they cannot happen.
The architecture is domain-agnostic from day one. Providers and domains are introduced by exposing capabilities through adapters. We are starting with one concrete implementation, not the full set of providers and domains the product will eventually cover. Adding a second provider within the same domain, a different domain entirely (flights, healthcare appointments, restaurants, local services), or a composable action flow should not require reworking the core scaffold.
The primary user is someone who wants to express intent in natural language and delegate the mechanics. They might say "Get me home," "Book me to the airport for 7am," or "Find me a dentist appointment next week." The user states the outcome they want, and the agent acts through the tools it has. The system serves either as a standalone concierge or as a capability inside a broader assistant, with rides, flights, appointments, and other tasks all reachable through the same scaffold.
The language model is the primary reasoner. It parses intent, decomposes compound requests, and decides which structured tool calls to make against the adapters. The durable plan serves as the model's scratchpad: a shared, inspectable workspace where the agent accumulates resolved context, candidate options, and pending actions as it works. The scaffold enforces the invariants (gates, logs, limits, isolation) that the model must not violate.
A plan is the durable record of one user intent. It holds the intent as the agent understood it, the resolved context gathered along the way, the current state (proposed, executed, in-progress, terminal), and any provider-side identifier once execution happens. One intent creates one plan, even when the intent involves multiple steps. Two intents create two plans, even in the same conversation. A plan ends in exactly one terminal state: completed, cancelled, or failed. Plans are the unit the scaffold tracks, the monitor watches, and the log indexes its entries against.
- Convert a natural-language user request into the right outcome: a confirmed, tracked, logged plan when the user wants action, or a quote or answer when the user is only seeking information. The user makes every consequential decision.
- Handle ambiguity by asking the user through structured prompts.
- Decompose multi-step or multi-domain requests into sub-tasks; each sub-task requires its own confirmation.
- Recover visibly when something goes wrong. Failures can be immediate (at execute time) or asynchronous (after commit). The user's original intent is preserved, the cause is explained, and a path forward is offered without requiring the user to re-enter the request. P0 assumes no post-commit failures occur; asynchronous failure recovery arrives with P1 active plan monitoring.
- Produce an audit trail sufficient to reconstruct any session after the fact.
- Treat multiple providers, multiple booking domains, and composable actions as a first-class property of the core system from the start.
- The system is not a real-time comparison aggregator. Price comparison and shopping are different products with different licensing constraints.
- The system does not replace native provider apps for users who already know exactly what they want.
- The system does not book anything silently. Every state-changing action requires an explicit, fresh user confirmation.
- The system does not try to substitute for provider error handling. Payment declines, inventory loss, dispatch issues, price adjustments, and similar provider-side failures are surfaced faithfully to the user with a recovery path; the scaffold does not re-implement the provider's own logic.
- The system does not support more than one active agent session per user at a time.
Smart model, dumb scaffold. The language model handles reasoning and conversation. It understands intent, decomposes compound requests, chooses which tools to call, and uses the durable plan as its scratchpad. The scaffold handles everything that must be correct every time: permission gates, confirmation state, duplicate prevention, logging, and recovery. The model operates within the space the scaffold allows.
-
Capability adapter interface. A clear contract exposing domain actions such as search, quote, book, track, cancel, and other composable operations. Each adapter declares which capabilities it supports and what contextual fields its domain requires, so the system branches on supported actions per adapter.
-
One working adapter. Implemented end to end as the first concrete proof of the model.
-
Confirmation gate. Every state-changing action splits into a proposal step and an execute step. The proposal is side-effect-free and presents the exact action to the user for review. Execution requires a fresh, single-use confirmation tied to that proposal. The system prevents replay, parameter substitution, and silent mutation between proposal and execution.
-
Action log. Append-only and durable, capturing every agent action, every user confirmation event, every provider interaction, and every state transition. The log is the source of truth; user-facing records derive from it. The log cannot be bypassed, overwritten, or disabled.
-
Quote re-validation before execution. Before any execute call, the system re-validates the quote (or equivalent time-bound commitment) against the provider. If the quote has expired or the terms have moved beyond tolerance, the existing proposal is invalidated and a fresh proposal is generated for new confirmation.
-
Intent resolution and request context. The scaffold captures what the agent needs to act: user identifier, payment method selection, named locations (home, work, saved addresses), time parsing (now, relative, absolute), and any domain-specific context fields the adapter declares. Context is session-scoped by default and separate from the P1 preference memory. Confirmations surface every resolved value, including any defaults applied without explicit user input, so users can catch and override wrong assumptions before committing.
-
Durable plan object as working memory. The plan is a durable, inspectable record that lives outside the conversation thread. It holds the agent's accumulated understanding of the user's request, the resolved context, candidate options, and pending actions. The scaffold can view, edit, and reason about the plan for monitoring, recovery, and scheduled execution; the agent uses it as a scratchpad while assembling structured tool calls. The plan is the bridge between the conversation and the adapters.
-
Single conversation channel and booking agent per user. Each user has exactly one conversation channel and one booking agent operating against that user's plan. Cross-provider and cross-domain conversations happen in that same channel, with the same agent coordinating. Continuity is owned at the scaffold level, independent of any individual provider.
-
Multi-user service with per-user isolation. The system serves many users concurrently. Each user reaches the agent through their own messaging channel (such as WhatsApp or another chat platform), and each user's plans, logs, context, and credentials are fully isolated from every other user's. No request on one user's channel can observe or affect another user's state.
-
Add more adapters across domains. Extend coverage by implementing additional adapters, both within the starting domain and in new domains. Each new adapter surfaces assumptions that were provider- or domain-specific in the first implementation and forces them into the shared abstraction.
-
Active plan monitoring. Monitoring is a per-user capability: each user has a single monitor that tracks all of that user's active plans. The monitor runs independently of any open conversation and outlives the chat session, the agent's turn, and the client connection.
The monitor tracks each plan's provider state with bounded staleness appropriate to the plan's phase. Every observed change writes to the action log before any user-facing notification fires, so the record of what the provider reported is preserved separately from the system's reaction to it.
Termination is explicit. A plan ends in exactly one terminal state: completed, cancelled, or failed. The monitor stops tracking a plan only when a terminal state is recorded. There is no implicit termination. If the provider stops responding and the outcome cannot be determined, the plan enters a recovery state that requires successful reconciliation or explicit user input to resolve, and remains in the log as unresolved until then.
Provider-reported events are surfaced to the user according to a default disposition (notify with action offered, notify only, or silent) that adapters declare per event class. Users can override the disposition for any event class.
-
Cross-provider fallback. When the primary provider fails with a known error class, the agent can propose the same intent routed to an alternative provider. Requires at least two adapters to exist. The provider switch requires explicit user confirmation; no automatic substitution.
-
Intent clarification as codified scaffold behavior. At P0 the LLM's own reasoning handles ambiguity. At P1 this becomes a first-class scaffold feature with structured clarification prompts and consistent formatting across domains.
-
Preference memory. Opt-in storage for stable preferences such as saved addresses, default payment method, accessibility needs, and domain-specific fields that adapters declare as preference-eligible. User-inspectable, user-editable, and user-deletable. Spending caps (per-session, per-day) live here as well; proposals that would exceed a cap require fresh confirmation before executing. Provider-held data is referenced as needed; it is not silently copied into long-lived product memory.
-
Cross-domain policy rules. Rules that span domains, such as a shared daily spend cap covering multiple booking types together.
-
Natural-language rules and preferences. Users can create, modify, and remove personal rules through conversation — spending limits ("don't let me spend more than $50 on rides today"), product preferences ("always recommend UberXL when going to the airport"), routing defaults ("prefer the cheapest option unless I say otherwise"), or any other standing instruction. The agent interprets the intent, persists the rule per-user, and applies it when relevant during future sessions. Users can ask "what are my rules?" to inspect, and can remove any rule the same way. Rules are user-owned: the agent cannot create or modify them without the user's explicit instruction.
At launch, the product succeeds when the following hold:
-
Log completeness. Every committed action has a complete log chain from initial intent to terminal state, reconstructible without session replay.
-
Gate integrity. The confirmation gate cannot be bypassed by any agent action, including pathological behavior.
-
Abstraction soundness. Exposing any booking domain through an adapter requires no changes to the scaffold, the agent interface, or the confirmation flow.
-
Postmortem sufficiency. A postmortem of any failed action can be conducted from the log alone.