Skip to content

About

End-to-end data science pipeline for WhatsApp conversation analysis — profiling, cleaning, sentiment analysis, embeddings, clustering

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

WhatsApp Interaction Analysis

Tests Quarto Publish Python 3.11+ R 4.5 Quarto

End-to-end data science pipeline for WhatsApp conversation analysis — profiling, cleaning, sentiment analysis, embeddings, clustering. Python + R, Quarto-rendered, CI/CD-deployed.

View the live site

About

Full data science pipeline for WhatsApp conversation analysis. The case study is a single export of ~92,000 messages spanning one year.

The project is designed to be reproducible — you can run the pipeline against new exports and integrate the results with the existing base.

Pipeline

Phase Stage Description
Preparation Data Discovery Initial exploration of the raw export
Data Profiling Systematic pattern investigation
Data Cleaning Invisible character removal, normalization
Data Wrangling Parsing, media linking, transcription
Feature Engineering 35+ derived variables
Model Features Multi-model sentiment analysis, embeddings
Analysis EDA Exploratory analysis by dimension (temporal, interaction, content)
Advanced Analysis Semantic clustering, PCA, MCA, N-grams, TF-IDF
Findings Consolidated insights Overview, dynamics, sentiment, themes, styles — written in prose

Structure

whatsapp-interaction-analysis/
│
├── index.qmd                         # Main document (overview)
├── .env.example                       # Configuration template
├── pyproject.toml                     # Packaging and dependencies
├── _quarto.yml                        # Main Quarto config
├── docs/BACKLOG.md                    # Prioritized backlog
├── docs/CHANGELOG.md                  # Project phase history
│
├── whatsapp/                          # Main package
│   ├── __init__.py                    # Version and metadata
│   ├── __main__.py                    # python -m whatsapp
│   ├── cli/                           # CLI (whatsapp-interaction)
│   │   ├── __init__.py                # Typer app + run command
│   │   ├── helpers.py                 # Shared helpers
│   │   ├── prepare.py                 # Commands: clean, wrangle, transcribe
│   │   ├── process.py                 # Commands: sentiment, embeddings
│   │   └── _status.py                 # Command: status
│   └── pipeline/                      # Pipeline modules
│       ├── config.py                  # Configuration (reads from .env)
│       ├── profiling.py               # Investigation functions
│       ├── cleaning.py                # Cleaning pipeline (7 stages)
│       ├── wrangling.py               # Wrangling pipeline (6 stages)
│       └── utils/                     # Utilities
│           ├── audit.py               # Audit system
│           ├── dataframe_helpers.py   # DataFrame helpers
│           ├── file_helpers.py        # File helpers
│           └── text_helpers.py        # Text helpers
│
├── scripts/                           # Standalone scripts
│   ├── transcribe_media.py            # Transcription via Groq/Whisper
│   ├── sentiment_*.py                 # Sentiment analysis (RoBERTa, DistilBERT, DeBERTa, ensemble)
│   ├── generate_embeddings*.py        # Embedding generation (mpnet, MiniLM, DistilUSE)
│   ├── compare_embeddings_models.py   # Embedding model comparison
│   ├── compare_embedding_dimensions.py # Embedding dimension comparison
│   ├── generate_timeline_chart.py     # "One Year in Data" timeline chart
│   └── generate_sample_data.py        # Synthetic dataset generator (for demo)
│
├── notebooks/                         # Quarto documents (see table below)
│
├── tests/                             # Unit tests (pytest)
│   ├── test_cleaning.py               # Cleaning pipeline tests
│   ├── test_wrangling.py              # Parsing and classification tests
│   └── test_cli.py                    # CLI tests
│
├── data/                              # Not versioned (personal data)
│   ├── raw/                           # Raw exports per period
│   ├── interim/                       # Intermediate files
│   ├── processed/                     # DataFrames per run
│   ├── external/                      # External context data
│   └── integrated/                    # Consolidated base
│
├── docs/
│   ├── SETUP-GUIDE.md                 # Installation guide
│   ├── INCREMENTAL-GUIDE.md           # Guide for new exports
│   └── data-dictionary.md             # Data dictionary
│
├── .github/workflows/
│   ├── tests.yml                      # Unit tests (Python 3.11/3.12)
│   └── publish.yml                    # Quarto site build (Python + R) → GitHub Pages
│
└── analysis/                          # Not versioned (outputs)

Notebooks

Preparation

# Notebook Description
00 Data Discovery Initial file exploration
00 Data Profiling Systematic investigation
01 Data Cleaning Cleaning and normalization
02 Data Wrangling Parsing, media, transcription
02.1 EDA — Data Wrangling Post-wrangling EDA
02.3 EDA — Content & Interaction Content analysis
03 External Context External data integration and relationship phases
04 Feature Engineering 35+ derived variables

Models (optional)

# Notebook Description
04 Model Features ML feature overview
04a Sentiment — RoBERTa Twitter-RoBERTa sentiment
04b Sentiment — DistilBERT DistilBERT sentiment
04c Sentiment — DeBERTa DeBERTa sentiment
04d Sentiment — Comparison Cross-model comparison
04e Sentiment — Ensemble 3-model ensemble
04f Embeddings — mpnet all-mpnet-base-v2
04g Embeddings — MiniLM all-MiniLM-L6-v2
04h Embeddings — DistilUSE distiluse-base-multilingual
04i Embeddings — Comparison Cross-model comparison

Analysis

# Notebook Description
05 EDA — Overview Exploratory analysis with context
05.1 EDA — Temporal Temporal patterns
05.2 EDA — Interaction Interaction dynamics
05.3 EDA — Content Content analysis
06 Advanced Analysis Semantic clustering, N-grams, TF-IDF

Findings

Synthesis layer in prose — consolidates pipeline findings into interpretive conclusions.

# Notebook Description
07 Findings — Overview Year overview + 4-axis synthesis
08 Findings — Dynamics Volume, rhythm and temporal patterns
09 Findings — Sentiment Dominant tone and emotional evolution
10 Findings — Themes Semantic clustering (k=10) and characterization
11 Findings — Styles P1 vs P2: vocabulary, emojis, punctuation

Lab

Experimental notebooks — explore alternative approaches and showcase the R + Python integration that defines the project's evolution. Rendered alongside the main pipeline via Quarto.

R + Python

Notebook Description
Visualization Gallery Creative WhatsApp viz ideas (ggplot2, gganimate, ggbump, ggwordcloud)
R Charts ggplot2 + plotly experiments
R Snippets Reusable ggplot patterns
Python ↔ R (reticulate) Mixing Python and R in the same document

Python variants

Notebook Description
Cleaning — Static Cleaning iteration (static output)
Cleaning — Hardened Cleaning with stricter cache settings
Wrangling — v1 Early wrangling iteration
Wrangling — Backup Wrangling backup snapshot

Quick Start

git clone https://github.com/mrlnlms/whatsapp-interaction-analysis.git
cd whatsapp-interaction-analysis

python3 -m venv .venv
source .venv/bin/activate

# Recommended: notebooks extra (covers viz, jupyter, scipy, sklearn)
pip install -e ".[notebooks]"

# Add ML on demand (transformers + torch, ~600 MB)
pip install -e ".[notebooks,ml]"

# Or pick exactly what you need:
# pip install -e ".[viz]"            # only visualization
# pip install -e ".[jupyter]"        # only Jupyter
# pip install -e ".[transcription]"  # only Groq transcription

cp .env.example .env
# Edit .env with your paths

quarto preview

Sample dataset

The repo ships with a synthetic dataset so you can test the pipeline without personal data:

# Configure .env to use sample data
echo "PROJECT_ROOT=$(pwd)" > .env
echo "DATA_FOLDER=sample" >> .env

# Run the preparation pipeline
whatsapp-interaction prepare clean
whatsapp-interaction prepare wrangle
whatsapp-interaction status

The sample dataset (200 messages, 7 days) is generated deterministically by scripts/generate_sample_data.py.

See the full Setup Guide.

CLI

pip install -e .

# Full pipeline
whatsapp-interaction run

# Preparation only (clean → wrangle → transcribe)
whatsapp-interaction prepare

# Individual steps
whatsapp-interaction prepare clean --steps u200e,anonymize
whatsapp-interaction process sentiment --model deberta

# Current state
whatsapp-interaction status

Audio transcription (optional)

# Add your API key to .env
echo "GROQ_API_KEY=your_key_here" >> .env

# Run the transcription script (~40 min for ~700 files)
python scripts/transcribe_media.py

# Re-render wrangling
quarto render notebooks/02-data-wrangling.qmd

The script auto-detects already-transcribed files and resumes where it stopped.

Tests

pip install pytest
pytest tests/ -v

149 tests covering the cleaning pipeline (cleaning.py), parsing/classification (wrangling.py) and the CLI. CI runs automatically via GitHub Actions on Python 3.11 and 3.12.

CI/CD

Two workflows run on every push to main:

  • tests.yml — pytest matrix on Python 3.11 and 3.12
  • publish.yml — Quarto build (Python 3.12 + R 4.5 with tidyverse, gganimate, ggbump, ggtext, ggwordcloud, plotly, reticulate) → deploy to GitHub Pages

The R toolchain enables the Lab notebooks above to render alongside the Python pipeline in a single CI run.

Tech Stack

Core: Python 3.11+, R 4.5 (CI), Quarto

Data: Pandas, NumPy, PyArrow

Visualization (Python): Matplotlib, Seaborn, Plotly, WordCloud

Visualization (R): tidyverse (ggplot2, dplyr, lubridate), gganimate, ggbump, ggtext, ggwordcloud, plotly

ML/Statistics: Scikit-learn, Prince (MCA), SciPy

NLP: Transformers/PyTorch (sentiment — RoBERTa, DistilBERT, DeBERTa), Sentence-Transformers (embeddings — mpnet, MiniLM, DistilUSE), Groq API/Whisper (transcription)

Interop: reticulate (Python in R documents)

Outputs

The pipeline generates the following files under data/processed/{export}/:

File Columns Description
messages.csv 8 Main analysis dataset
messages.parquet 8 Same content, ~3x smaller
messages_full.csv 17 Full version for debugging
chat_complete.txt — Chat with transcriptions
corpus_*.txt — NLP-ready text

Documentation

Privacy

Data folders (data/ and analysis/) are not versioned because they contain personal information.


Built by @mrlnlms

About

End-to-end data science pipeline for WhatsApp conversation analysis — profiling, cleaning, sentiment analysis, embeddings, clustering

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages