Skip to content

About

Open-source, agent-native qualitative research platform. Code interviews with reviewable AI suggestions, confirmed-evidence states, and a full audit trail. A transparent alternative to NVivo, ATLAS.ti, and Dovetail. Apache-2.0.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

OpenVerbatim

Open-source, agent-native qualitative research platform. Code interviews with AI help, keep every AI suggestion under human review, and trace every answer back to the exact words a participant said. A self-hostable alternative to NVivo, ATLAS.ti, and Dovetail.

License Node TypeScript

Try it without installing anything: sandbox.openverbatim.com runs the full review loop in your browser on generated demo data. No account, no upload, nothing leaves your machine.

Why we built this

Two things pushed us to start this project.

The first is cost and control. NVivo and ATLAS.ti are now owned by the same company, they are expensive, and a graduate student on a stipend often cannot justify the license. The data you put into them, in most cases, lives on someone else's terms.

The second is trust. Interview data is some of the most sensitive material a researcher handles: named people, health details, things said in confidence. Handing that to a closed cloud tool, or letting an AI quietly rewrite your codes, is a hard thing to defend to an ethics board. Recent journal and IRB guidance on disclosing AI use has raised the bar for what researchers need to be able to account for.

OpenVerbatim answers both. It is Apache-2.0, so you can read the code, self-host it, and keep your data on machines you control. And it runs on one rule: AI output stays a suggestion until a human accepts it, and every state change is written to an append-only audit trail. You can always show what the AI proposed, what a person decided, and when.

What it does

You import interview audio or transcripts, get AI-assisted open coding that you accept, edit, or reject one keystroke at a time, group confirmed codes into themes, and then ask questions of the dataset with answers that cite the exact source lines. Nothing becomes a finding until a person has signed off on the evidence behind it.

Area Capability
Audio ingest Upload interview audio, store it locally, run preview plus authoritative transcription with diarization.
Streaming transcription Chunked preview transcription streams progress over SSE; a longer authoritative transcript can replace and re-anchor evidence.
Agent open coding Suggestions carry quotes, rationale, confidence, and grounding checks that drop fabricated or unlocatable quotes.
Three-pane workspace Source list, transcript, and review/codebook panes with accept/edit/reject keyboard flow.
Autonomy dial Per-project review, cruise, and full-auto modes route low-confidence or new suggestions into an exception queue.
Theme clustering Confirmed codes group into themes while keeping their membership intact.
Ask your data Questions run against confirmed evidence only; answers cite verbatim quotes with timestamp jumps.
Audit and provenance Every pipeline, review, processor, and answer action is paired with an append-only audit event.
Consent withdrawal Withdrawing a participant cascades a real deletion across transcripts, codes, embeddings, and shared excerpts, and records that the deletion happened.
Browser sandbox A zero-backend demo runs against generated local data using SQLite WASM.

Try it

Quick start

Local development

pnpm install
cp .env.example .env
pnpm dev

The API defaults to http://localhost:8787, the web app runs on Vite, and the sandbox runs as a separate Vite app. For live server-backed workflows, start Postgres first and set DATABASE_URL.

With Postgres

docker compose -f infra/docker-compose.dev.yml up -d postgres
export DATABASE_URL=postgres://openverbatim:openverbatim@127.0.0.1:55433/openverbatim
export TEST_DATABASE_URL=postgres://openverbatim:openverbatim@127.0.0.1:55433/openverbatim_test
pnpm install
pnpm typecheck
pnpm test

The compose file uses pgvector/pgvector:pg16 because the retrieval layer stores vector embeddings in Postgres.

Bring your own AI keys: transcription and coding call out to providers you configure (OpenRouter, Deepgram, AssemblyAI). Nothing is hardcoded to a vendor, and the self-hosted build has no paid feature wall.

Configuration

Copy .env.example to .env and fill only the providers you plan to use. OpenRouter values are plain model IDs; the default anthropic/claude-opus-4.8 can be swapped for any model your account can reach.

Variable Default Purpose
DATABASE_URL postgres://openverbatim:openverbatim@127.0.0.1:55433/openverbatim Main Postgres database URL.
TEST_DATABASE_URL postgres://openverbatim:openverbatim@127.0.0.1:55433/openverbatim_test Isolated test database; tests create it when needed.
PORT 8787 API server port.
OPENROUTER_API_KEY unset BYOK key for OpenRouter-backed coding, clustering, ask, embeddings, and preview ASR.
ASSEMBLYAI_API_KEY unset BYOK key for the AssemblyAI authoritative transcript adapter.
DEEPGRAM_API_KEY unset BYOK key for the Deepgram authoritative transcript adapter.
OV_STORAGE_DIR apps/server/.openverbatim-storage Local storage root for uploaded source files.
OV_ASR_PREVIEW_MODEL openai/whisper-large-v3-turbo OpenRouter audio preview model.
OV_ASR_AUTHORITATIVE deepgram Authoritative ASR provider: deepgram or assemblyai.
OV_CODING_MODEL anthropic/claude-opus-4.8 Primary OpenRouter model for coding suggestions.
OV_CLUSTERING_MODEL anthropic/claude-opus-4.8 Primary OpenRouter model for theme clustering.
OV_ASK_MODEL anthropic/claude-opus-4.8 Primary OpenRouter model for ask-your-data answers.
OV_EMBEDDING_PROVIDER hash Embedding provider: hash for local deterministic tests or openrouter.
OV_AUTONOMY_CONFIDENCE_THRESHOLD 0.75 Confidence threshold used by the autonomy policy engine.
VITE_OV_API_BASE http://localhost:8787 Web app API base URL.
VITE_OV_LIVE unset Set to 1 to use live API data in the web app.

The full list of tuning variables (ASR budgets, retrieval thresholds, re-anchoring tolerances) is in .env.example.

Architecture

OpenVerbatim is a TypeScript monorepo with three apps and two shared packages:

apps/web        React + Vite researcher workspace
apps/server     Hono API, SSE stream, queue worker, provider adapters
apps/sandbox    Browser-only demo with SQLite WASM
packages/db     Postgres schema, migrations, data access, audit events
packages/shared Shared types, paths, and event contracts

Runtime shape:

          React workspace             Browser sandbox
                |                            |
          REST + SSE                         | local SQLite WASM
                v                            v
        Hono API server  <---------->  Postgres + pgvector
                |
          queue worker
                |
   ASR / LLM / embedding adapters

The API and worker share one Postgres database. Pipeline jobs are claimed from the database, external processor calls are checked against project policy, and user-visible updates stream over SSE.

Roadmap

  • Researcher workspace: project setup, source list, transcript review, keyboard review flow.
  • API: project, source upload, pipeline run, adjudication, autonomy, audit, theme, and ask endpoints.
  • Worker: text and audio preview paths, authoritative ASR, coding windows, autonomy policy, theme clustering.
  • Postgres schema, migrations, audit events, pgvector retrieval, isolated test database.
  • Consent withdrawal with cascade deletion and audit tombstoning.
  • Browser sandbox with generated demo data.
  • In progress: REFI-QDA export and import for interchange with other CAQDAS tools.
  • Planned: inter-coder agreement report export.
  • Planned: blind-mode review that hides agent suggestions before reveal.

Project status and sustainability

This is early software under active development, built by a team rather than a single maintainer. It is Apache-2.0, so the project can be forked and kept alive independently of us. Development is funded by an optional hosted version at openverbatim.com; the self-hosted build stays complete and free, with no feature held back for the paid tier.

If you are evaluating it for a study, treat the current release as a working beta and check the roadmap for what is not built yet.

Contributing

See CONTRIBUTING.md.

License

Apache-2.0. See LICENSE.

About

Open-source, agent-native qualitative research platform. Code interviews with reviewable AI suggestions, confirmed-evidence states, and a full audit trail. A transparent alternative to NVivo, ATLAS.ti, and Dovetail. Apache-2.0.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages