Skip to content
Anouâr Gadermann edited this page Aug 25, 2026 · 12 revisions
CapsQual_logo

CapsQual User Manual

Table of Contents

1. Introduction

2. Projects

3. Importing Subtitle Files

4. Basic Formatting Functions

5. Automatic Editing Functions

6. Audio Functions

7. Exporting Transcripts

8. Settings

9. Keyboard Shortcuts

10. Troubleshooting

1. Introduction

1.1 Overview

CapsQual is a specialized transcription workstation designed for converting subtitle files into publishable interview transcripts using GAT2 (Gesprächsanalytisches Transkriptionssystem 2) conventions.

1.2 Key Features

  • Quick Speaker Assignment: Rapidly assign speakers to subtitle segments using keyboard shortcuts
  • Basic Editing: Split, merge, and edit transcript segments
  • Transcription Symbols: Insert conversation analysis symbols for pauses, breathing, and annotations (GAT2, TiQ and Dresing & Pehl)
  • Audio Synchronization: Link audio files to transcripts with auto-sync functionality and a waveform viewer
  • Export Transcripts: Generate formatted transcripts in HTML, DOCX, TXT or SRT

2. Projects

2.1 Basic Project File Management

  • Start a new transcription project by selecting File → New Project or pressing Ctrl+N.
  • Open saved projects with File → Open Project or Ctrl+O. CapsQual uses the .capsqual file format.
  • Recently opened projects may be opened through the "Open Recent" submenu.
  • Save your work with File → Save Project (Ctrl+S) or Save Project As to create a new file.

2.2 Project Memos

Add project notes and descriptions via Edit → Project Memo. This information can be included in transcript exports.

3. Importing Subtitle Files

3.1 Supported File Formats

  • SRT (.srt): Standard subtitle format with timestamps
  • VTT (.vtt): WebVTT subtitle format with timestamps; may include speaker tags and bold/italic/underline formatting (see 3.4)
  • JSON (.json): Various JSON formats including token-based and segment-based
  • Text (.txt): Plain text files (one block per line)
  • TSV (.tsv): Tab-separated values with start/end times and text

3.2 Import Methods

Use File → Import → Subtitles to load subtitle files. When importing audio files, CapsQual automatically searches for matching subtitle files in the same directory and offers to import them.

Files can also be imported by dragging and dropping them directly onto the CapsQual window:

  • Subtitle files (.srt, .vtt, .txt, .json, .tsv) replace the current transcript
  • Audio files load into the current project
  • Project files (.capsqual, .capsgat) open as projects

3.3 JSON Import Options

When importing JSON files with token data, you can choose from three import methods:

  • One Continuous Block: Import all text as a single segment
  • Tokens as Separate Blocks: Create individual blocks for each token
  • Auto-segment: Automatically detect pauses and create segments accordingly

3.4 VTT Speaker Diarization and Formatting

WebVTT files that use <v SpeakerName> voice tags (as produced by some captioning and transcription tools) have their speaker names imported automatically. Formatting tags such as <b>, <i> and <u> are converted into CapsQual's formatting markers, and <c.class> span tags are removed.

4. Basic Formatting Functions

4.1 Speaker Assignment

Assign speakers using the number keys 1-4 (configurable up to 8 speakers). Unassign with U.

4.2 Navigation

Keyboard Shortcuts:

  • Next block: N or →
  • Previous block: P or ←
  • Jump to unassigned: Click on block in "Unassigned Blocks" list
  • Full shortcut reference: Press [F1] (or Help → Shortcuts) to open the searchable keyboard-shortcuts dialog

4.3 Editing Functions

  • Split Block: Space - Opens split dialog to divide current segment
  • Merge Blocks: Delete - Combine current block with next block
  • Edit Content: E - Open text editor for current block
  • Insert Empty Line: Enter - Add blank line for formatting

4.4 Transcript Symbols

  • Access the symbols dialog with * or the Symbols button.
  • Switch between symbols for GAT2, Dresing & Pehl, TiQ and custom symbol tabs using the [Tab] key.
  • Custom symbols can be defined in the custom symbol tab. These can also be saved and restored as .JSON files.
  • Custom symbols can also be managed from the menu bar (Edit → Custom Symbols).

5. Automatic Editing Functions

5.1 Insert pauses for gaps between segments

  • Open the dialog from Edit → Modify transcript → Insert pauses for gaps
  • Choose whether to use GAT2, Dresing & Pehl and TiQ style pause symbols.
  • Choose whether pauses should be inserted as separate segments, at the start of segments or both – depending on a custom threshold.
  • Set a threshold for pause notation (e.g., to ignore all pauses under 0.2 seconds).
  • Set a threshold for numeric pause notation (e.g., to use numerals for all pauses larger than 2 seconds).

5.2 Strip punctuation

  • To strip the entire transcript of punctuation marks, use Edit → Modify transcript → Strip punctuation.
  • Confirm the dialog to proceed. Note that all punctuation, including from transcription symbols will be removed.
  • Formatting markers (bold, underline, italic) will not be affected.

5.3 Convert to lowercase

  • To change Latin-based transcript to lowercase, use Edit → Modify transcript → Convert to lowercase.
  • Confirm the prompt to proceed. Note that the text will be converted to all lowercase.
  • Non-Latin transcript text as well as textual formatting markers will not be affected.

6. Audio Functions

6.1 Importing Audio

Load audio files via File → Import → Audio File (or drag and drop them onto the window). Supported formats: MP3, WAV, OGG, M4A, FLAC, AAC, WMA.

6.2 Automatic Subtitle Detection

When importing audio, CapsQual automatically searches the directory for matching subtitle files (SRT, VTT, JSON, TXT, TSV) and offers to import them.

6.3 Audio Controls

Playback Controls:

  • Play/Pause: [End] or click ▶ or ⏸ button
  • Rewind 5s: [PgUp] or click ⏪ button
  • Fast Forward 5s: [PgDn] or click ⏩ button
  • Jump to Time: Click on the time display or press [Ctrl]+[J] to jump to a time in the audio file.
  • Play from segment: Press the "Play from segment" button or [Shift]+[Enter] to start playback from the current segment.
  • Adjust speed: Use the mouse wheel on the knob, the buttons or the + and - keys to adjust playback speed (VLC Player needs to be installed for this feature to work!)

6.4 Additional Audio Features

  • Auto-sync to Audio: Automatically highlights the current transcript block during playback (Using audio controls [PgUp/PgDwn] instead of scrolling is recommended)
  • Autopause: Automatically pauses audio when opening editing dialogs
  • Progress Tracking: Visual progress bar with time display

6.5 Waveform Viewer

The waveform viewer displays the audio waveform together with the current segment's start (left) and end (right) markers. It is an alternative to editing timestamps manually.

  • Adjust a boundary by dragging: drag the start or end marker to a new position. The whole drag is treated as a single undo step.
  • Nudge the start marker: [Ctrl]+[Left] / [Ctrl]+[Right]
  • Nudge the end marker: [Ctrl]+[Shift]+[Left] / [Ctrl]+[Shift]+[Right]
  • Set a marker to the current playback position: [Ctrl]+[Shift]+[Alt]+[Left] (start marker) or [Ctrl]+[Shift]+[Alt]+[Right] (end marker)
  • Zoom: use the + / − buttons on the right edge of the viewer.

The keyboard shortcuts work from anywhere in the window (no need to click the waveform first). Holding a key repeats the nudge, and the whole burst counts as one undo step.

7. Exporting Transcripts

7.1 Export Formats

Generate transcripts in these formats:

  • HTML: Formatted with CSS styling
  • Word document: A publication-ready .DOC file
  • Plain Text: Simple text format for maximum copy-paste compatibility
  • Subtitle file: An .SRT file containing time stamps and speaker diarization

7.2 Transcript Conventions

Transcript exports can be customized to better suit one of several transcription systems:

  • GAT2-Transcription: More commonly used in conversation analysis; each segment takes up a line, a mono-spaced font is used, speaker-labels and line numbering follow GAT2 guidelines.
  • Dresing & Pehl: More commonly used for semantic analysis; segments are arranged by speaker turns.
  • TiQ (Talk in Qualitative Research): Segments are arranged by speaker turns, often used for reconstructive research (especially for group discussions).

Please note that choosing transcription conventions does not alter the transcript content, but only the basic transcript formatting. The choice should be based on whichever format best fits the conventions used during transcription.

7.3 Export Options

Customize your export with these options:

  • Line-Wrapping: When enabled, line breaks will be implemented in the exported transcript. Enter the amount of characters, after which line breaks should occur. Enabling "Force character-based wrapping" results in line breaks occurring regardless of words, whereas disabling this checkbox results in line breaks after spaces.
  • Include Timestamps: Add timecodes to transcript (when available); these can follow different formats. For custom timestamp formats, the letter-combinations HH, mm, ss and xx can be used to represent hours, minutes, seconds and tenths of seconds respectively. Any other symbols, letters, numbers, characters etc. can be used alongside these to make up a custom timestamp formula.
  • Include Diarization (Speaker Labels): Choose whether to include speaker labels (only applicable for subtitle exports)
  • Add empty line after every speaker turn: An empty line is added after every turn. Empty lines are handled differently for different conventions. For TiQ, empty lines are numbered, whereas for GAT2 and Dresing & Pehl they are not.
  • Project Title: Include project name as header
  • Project Memo: Include project description
  • Audio File Info: Include source audio file name

7.4 Export Preview

Use the preview feature to review your transcript formatting before exporting. The preview shows an approximation of how the final transcript will appear.

8. Settings

Open the settings dialog via Edit → Settings... to customize the following:

  • Theme: switch between Light and Dark mode with the toggle switch. The change is applied immediately and remembered for future sessions.
  • Default path: choose whether file dialogs start in the system default location or a custom directory of your choice.
  • Text Display Font: select the font used for the transcript display.
  • Optimize for CJK: enable double spaces for overlap indentation in CJK (Chinese/Japanese/Korean) text.

9. Keyboard Shortcuts

Press [F1] (or Help → Shortcuts) to open the built-in searchable shortcut reference at any time. The dialog groups shortcuts by category and lets you type to filter the list, so it stays manageable even as new shortcuts are added.

The full shortcut reference is also reproduced below.

Navigation

Shortcut Action
P / Left Arrow Previous block
N / Right Arrow Next block

Assigning Speakers

Shortcut Action
1-4 Assign speakers A-D
U Unassign current block

Editing

Shortcut Action
Space Split current block
Delete Merge with next block
E / F2 Edit segment content
T Edit segment timestamp
Enter Insert empty line
Ctrl+Del Remove overlap

Transcription Symbols

Shortcut Action
* Open symbols dialog
. Insert micropause (with placement)
h Insert short inhale (with placement)
H Insert short exhale (with placement)

Audio Controls

Shortcut Action
End Play/Pause audio
PgUp Rewind 5 seconds
PgDn Fast forward 5 seconds
Ctrl+J Jump to Time
Shift+L Jump to Current Audio Location
Ctrl+L Toggle Auto-sync to Audio
Shift+Enter Play from current segment
- Lower Playback Speed
+ Speed up Playback

Waveform Viewer

Shortcut Action
Ctrl+Left/Right Move left marker
Ctrl+Shift+Left/Right Move right marker
Ctrl+Shift+Alt+Left/Right Set right/left marker to audio position

Search Functions

Shortcut Action
Ctrl+F Open Search Dialog
F3 Find Next
Shift+F3 Find Previous

File Operations

Shortcut Action
Ctrl+N New Project
Ctrl+O Open Project
Ctrl+S Save Project
Ctrl+Return Export Transcript

Help

Shortcut Action
F1 Show Shortcuts (this dialog)
Ctrl+F1 Open Online Manual (requires internet)

10. Troubleshooting

This section collects common problems and their solutions. If your issue is not listed here, please open an issue with your operating system, the CapsQual version, and a short description of what went wrong.

10.1 Audio not loading

Symptom: Importing an audio file does nothing, or playback buttons are disabled.

  • VLC not installed. CapsQual's audio features require VLC media player. Install it and restart CapsQual.
  • VLC installed but audio still not loading. Try reinstalling VLC. On macOS, install it via Homebrew:
    brew install vlc
    
  • Waveform does not render. CapsQual uses soundfile to read audio for the waveform. If you run CapsQual from source, make sure the dependencies are installed (pip install -r requirements.txt). The file also needs to be in a supported format (MP3, WAV, OGG, M4A, FLAC, AAC, WMA).

10.2 Audio stopped working

Symptom: Audio worked earlier, but playback suddenly fails (the audio player can crash occasionally).

Save your work, then reopen the file via File → Open Recent. This reloads the project and starts a fresh audio player instance.

10.3 Audio files don't get transcribed when opened

CapsQual is not automatic speech recognition (ASR) software and does not transcribe audio on its own. It is designed to turn the output of transcription tools (such as Whisper AI or CapsWriter) into qualitative research transcripts. Load the subtitle file produced by such a tool (see Importing Subtitle Files) and attach the audio for synchronization.

10.4 Speaker assignment via number keys doesn't work

The number keys (1–4) assign speakers only when the transcript has focus. If a text field (e.g. the search box or a speaker name field) has focus, the keys type into that field instead. Click on the transcript area first, then press the number key.

10.5 Custom symbols don't appear on another computer

Custom symbols are stored per user in the operating system's application-data directory (for example %APPDATA%\CapsQual on Windows), not inside the program folder. They do not travel with a project or an installation. Use Export Custom Symbols / Import Custom Symbols (Edit → Custom Symbols) to transfer your symbols between computers.

10.6 Auto-sync to audio doesn't highlight the transcript

Auto-sync requires the transcript to contain timestamps, and it needs to be enabled with [Ctrl]+[L] or the Auto-sync checkbox. It cannot follow playback when the transcript has no time information.

10.7 DOCX export produces a TXT file instead

When CapsQual runs from source and python-docx is not installed, the DOCX export silently falls back to plain text. Install the dependencies (pip install -r requirements.txt) to enable Word document export. Packaged releases include python-docx, so this usually only affects source installations.

10.8 Transcript looks messy / indentation doesn't align

Make sure an equidistant (monospaced) font is used, such as Consolas or Courier New. GAT2 and TiQ transcripts default to Courier New. Proportional fonts break column alignment of timestamps, speaker labels and line numbers.

10.9 Crash after deleting all custom symbols, or VTT shows "No content loaded"

These were bugs in versions before 1.6.3. Updating to the latest release fixes:

  • Crashes when reopening the symbol dialog after deleting all custom symbols.
  • VTT files with MM:SS.mmm timestamps (no hours component) showing "No content loaded".

10.10 CapsQual crashed — what now?

  • Save your work frequently.
  • Reopen the file via File → Open Recent — this also reloads the audio player.
  • If the crash repeats, open an issue and include your operating system, the CapsQual version, and the exact steps that caused the crash.

10.11 Windows shows "Windows protected your PC" / unknown publisher

CapsQual is not code-signed. When Windows blocks a downloaded installer or executable, click More info, then Run anyway. On macOS, right-click (or Ctrl-click) the app, select Open, then Open again to bypass the unidentified-developer warning.

Clone this wiki locally