Software engineer based in New York, working on distributed data systems.
My day-to-day work centers on Spark, Hadoop, Java, and Python: building and supporting large-scale ETL and processing workloads across distributed storage and compute environments.
I am especially interested in:
- Modern data platforms, including object storage and Iceberg-style lakehouse architectures.
- Software quality at scale: static analysis, CI/CD, SonarQube, and developer tooling for JVM codebases.
- Reliable agentic engineering workflows: grounded code-review assistance, evaluation, guardrails, and human review.
- Open-source software and sustainable developer communities.
| Area | Technologies |
|---|---|
| Languages | |
| Storage | |
| Processing | |
| Streaming | |
| Tooling |
| Repository | About |
|---|---|
| apache/airflow | Apache Airflow - A platform to programmatically author, schedule, and monitor workflows |
| apache/solr | Apache Solr open-source search software |
| apache/fineract-backoffice-ui | Angular back-office UI for Apache Fineract, the open-source core banking platform |
| apache/datafusion | Apache DataFusion SQL Query Engine |
| gchq/sleeper | A cloud-native, serverless, scalable, cheap key-value store |
| apache/sedona | A cluster computing framework for processing large-scale geospatial data |
| apache/arrow-rs | Official Rust implementation of Apache Arrow |
| apache/security-dash | Apache security |

