Open-source, agent-native qualitative research platform. Code interviews with AI help, keep every AI suggestion under human review, and trace every answer back to the exact words a participant said. A self-hostable alternative to NVivo, ATLAS.ti, and Dovetail.
Try it without installing anything: sandbox.openverbatim.com runs the full review loop in your browser on generated demo data. No account, no upload, nothing leaves your machine.
Two things pushed us to start this project.
The first is cost and control. NVivo and ATLAS.ti are now owned by the same company, they are expensive, and a graduate student on a stipend often cannot justify the license. The data you put into them, in most cases, lives on someone else's terms.
The second is trust. Interview data is some of the most sensitive material a researcher handles: named people, health details, things said in confidence. Handing that to a closed cloud tool, or letting an AI quietly rewrite your codes, is a hard thing to defend to an ethics board. Recent journal and IRB guidance on disclosing AI use has raised the bar for what researchers need to be able to account for.
OpenVerbatim answers both. It is Apache-2.0, so you can read the code, self-host it, and keep your data on machines you control. And it runs on one rule: AI output stays a suggestion until a human accepts it, and every state change is written to an append-only audit trail. You can always show what the AI proposed, what a person decided, and when.
You import interview audio or transcripts, get AI-assisted open coding that you accept, edit, or reject one keystroke at a time, group confirmed codes into themes, and then ask questions of the dataset with answers that cite the exact source lines. Nothing becomes a finding until a person has signed off on the evidence behind it.
| Area | Capability |
|---|---|
| Audio ingest | Upload interview audio, store it locally, run preview plus authoritative transcription with diarization. |
| Streaming transcription | Chunked preview transcription streams progress over SSE; a longer authoritative transcript can replace and re-anchor evidence. |
| Agent open coding | Suggestions carry quotes, rationale, confidence, and grounding checks that drop fabricated or unlocatable quotes. |
| Three-pane workspace | Source list, transcript, and review/codebook panes with accept/edit/reject keyboard flow. |
| Autonomy dial | Per-project review, cruise, and full-auto modes route low-confidence or new suggestions into an exception queue. |
| Theme clustering | Confirmed codes group into themes while keeping their membership intact. |
| Ask your data | Questions run against confirmed evidence only; answers cite verbatim quotes with timestamp jumps. |
| Audit and provenance | Every pipeline, review, processor, and answer action is paired with an append-only audit event. |
| Consent withdrawal | Withdrawing a participant cascades a real deletion across transcripts, codes, embeddings, and shared excerpts, and records that the deletion happened. |
| Browser sandbox | A zero-backend demo runs against generated local data using SQLite WASM. |
- Hosted sandbox (no install): sandbox.openverbatim.com
- Self-host: see the quick start below.
pnpm install
cp .env.example .env
pnpm devThe API defaults to http://localhost:8787, the web app runs on Vite, and the sandbox runs as a separate Vite app. For live server-backed workflows, start Postgres first and set DATABASE_URL.
docker compose -f infra/docker-compose.dev.yml up -d postgres
export DATABASE_URL=postgres://openverbatim:openverbatim@127.0.0.1:55433/openverbatim
export TEST_DATABASE_URL=postgres://openverbatim:openverbatim@127.0.0.1:55433/openverbatim_test
pnpm install
pnpm typecheck
pnpm testThe compose file uses pgvector/pgvector:pg16 because the retrieval layer stores vector embeddings in Postgres.
Bring your own AI keys: transcription and coding call out to providers you configure (OpenRouter, Deepgram, AssemblyAI). Nothing is hardcoded to a vendor, and the self-hosted build has no paid feature wall.
Copy .env.example to .env and fill only the providers you plan to use. OpenRouter values are plain model IDs; the default anthropic/claude-opus-4.8 can be swapped for any model your account can reach.
| Variable | Default | Purpose |
|---|---|---|
DATABASE_URL |
postgres://openverbatim:openverbatim@127.0.0.1:55433/openverbatim |
Main Postgres database URL. |
TEST_DATABASE_URL |
postgres://openverbatim:openverbatim@127.0.0.1:55433/openverbatim_test |
Isolated test database; tests create it when needed. |
PORT |
8787 |
API server port. |
OPENROUTER_API_KEY |
unset | BYOK key for OpenRouter-backed coding, clustering, ask, embeddings, and preview ASR. |
ASSEMBLYAI_API_KEY |
unset | BYOK key for the AssemblyAI authoritative transcript adapter. |
DEEPGRAM_API_KEY |
unset | BYOK key for the Deepgram authoritative transcript adapter. |
OV_STORAGE_DIR |
apps/server/.openverbatim-storage |
Local storage root for uploaded source files. |
OV_ASR_PREVIEW_MODEL |
openai/whisper-large-v3-turbo |
OpenRouter audio preview model. |
OV_ASR_AUTHORITATIVE |
deepgram |
Authoritative ASR provider: deepgram or assemblyai. |
OV_CODING_MODEL |
anthropic/claude-opus-4.8 |
Primary OpenRouter model for coding suggestions. |
OV_CLUSTERING_MODEL |
anthropic/claude-opus-4.8 |
Primary OpenRouter model for theme clustering. |
OV_ASK_MODEL |
anthropic/claude-opus-4.8 |
Primary OpenRouter model for ask-your-data answers. |
OV_EMBEDDING_PROVIDER |
hash |
Embedding provider: hash for local deterministic tests or openrouter. |
OV_AUTONOMY_CONFIDENCE_THRESHOLD |
0.75 |
Confidence threshold used by the autonomy policy engine. |
VITE_OV_API_BASE |
http://localhost:8787 |
Web app API base URL. |
VITE_OV_LIVE |
unset | Set to 1 to use live API data in the web app. |
The full list of tuning variables (ASR budgets, retrieval thresholds, re-anchoring tolerances) is in .env.example.
OpenVerbatim is a TypeScript monorepo with three apps and two shared packages:
apps/web React + Vite researcher workspace
apps/server Hono API, SSE stream, queue worker, provider adapters
apps/sandbox Browser-only demo with SQLite WASM
packages/db Postgres schema, migrations, data access, audit events
packages/shared Shared types, paths, and event contracts
Runtime shape:
React workspace Browser sandbox
| |
REST + SSE | local SQLite WASM
v v
Hono API server <----------> Postgres + pgvector
|
queue worker
|
ASR / LLM / embedding adapters
The API and worker share one Postgres database. Pipeline jobs are claimed from the database, external processor calls are checked against project policy, and user-visible updates stream over SSE.
- Researcher workspace: project setup, source list, transcript review, keyboard review flow.
- API: project, source upload, pipeline run, adjudication, autonomy, audit, theme, and ask endpoints.
- Worker: text and audio preview paths, authoritative ASR, coding windows, autonomy policy, theme clustering.
- Postgres schema, migrations, audit events, pgvector retrieval, isolated test database.
- Consent withdrawal with cascade deletion and audit tombstoning.
- Browser sandbox with generated demo data.
- In progress: REFI-QDA export and import for interchange with other CAQDAS tools.
- Planned: inter-coder agreement report export.
- Planned: blind-mode review that hides agent suggestions before reveal.
This is early software under active development, built by a team rather than a single maintainer. It is Apache-2.0, so the project can be forked and kept alive independently of us. Development is funded by an optional hosted version at openverbatim.com; the self-hosted build stays complete and free, with no feature held back for the paid tier.
If you are evaluating it for a study, treat the current release as a working beta and check the roadmap for what is not built yet.
See CONTRIBUTING.md.
Apache-2.0. See LICENSE.