Aaron O'Brien / bioinformatics + software

Aaron O'Brien

I build the software that turns long sequencing reads into taxonomy you can trust, ITSxRust, EMITS and SSUplex, written in Rust and used in fungal and environmental metabarcoding workflows. Day to day I run the Nanopore amplicon work at the Centro de Biotecnología de Sistemas in Santiago, from library handoff to the dashboard the biologist actually opens.

Centro de Biotecnología de Sistemas, Universidad Andrés Bello Santiago, Chile
amplicon design
A long read split into ribosomal regions The ITS region sits between conserved ribosomal flanks. Switching the primer pair changes how much flanking sequence a read carries, and whether all four profile-HMM anchors can be matched.
4/4anchors matched

Extraction rates and timings from O'Brien et al. 2026, Methods in Ecology and Evolution.

Tools

Open source

ITSxRust

ITS extraction for long reads

Finds ITS1, 5.8S and ITS2 inside Nanopore and PacBio amplicons by chaining four profile-HMM anchors, and falls back to two-anchor pairs when a read is missing a flank, so every region it returns still has a model boundary on both sides rather than a read edge. Failures come back as structured codes and a per-sample QC summary instead of a silently shorter FASTA.

Rustprofile-HMMBiocondanf-core/ampliseq
ITSxRust75.3%
ITSx69.9%
ITSxpress v241.4%

Full ITS recovered from a 54,659-read Nanopore library, at 4.6× the speed of ITSx.

EMITS

abundance estimation

Reads that map equally well to several UNITE references are the reason species-level abundance tables lie. EMITS runs expectation-maximization over minimap2 alignments to split those reads probabilistically and aggregate across redundant accessions, with presets for Nanopore and PacBio chemistries.

Rustexpectation-maximizationminimap2UNITE

SSUplex

small-subunit extraction

Pulls small-subunit rRNA out of environmental reads on either strand and sorts what it finds by origin, which matters when host or organellar sequence would otherwise swamp the target community.

RusteDNA16S / 18S

Built for the lab

CSB-LIMS

sample tracking

Where a sample came from, what was done to it and which run it ended up in — with role-based access so the people doing the bench work can enter their own data.

FastAPIReactPostgreSQL

BioBalance

bioleaching campaigns

Metal balances and recovery curves across minicolumn campaigns, deployed on site so results are current rather than reconstructed from spreadsheets at the end.

FastAPIPostgreSQLDocker

Amplicon pipelines

16S and ITS, end to end

Snakemake and Nextflow workflows for the centre's Nanopore runs, with PICRUSt2 functional profiling and interactive D3 reports for vineyard soil, bioleaching and acid mine drainage, compost time series and food-safety surveillance.

SnakemakeNextflowQIIME2D3.jsSLURM

Work

ten years, industry and academia

M.S. Bioinformatics, University of Maine (2020–2024), taken alongside the Kraken and UVA roles. B.S. Computer Science, Florida Southern College (2014–2016).

Papers

Get in touch

Happy to talk about long-read metabarcoding, pipeline engineering, or collaborations where the sequencing is done and the analysis needs to hold up.