Comparative analysis of transcriptome-derived T-cell receptor repertoires across the human respiratory tract reveals compartment-associated variation and similar CDR3 sequence features.

Published: 18 September 2026| Version 3 | DOI: 10.17632/mrchhxsgdt.3
Contributor:
Alex Kayongo

Description

Adaptive immune responses are distributed across anatomically distinct compartments of the human respiratory tract, yet comparative characterization of transcriptome-derived T-cell receptor (TCR) repertoires across these compartments remains limited. Here, we reconstructed transcriptome-derived T-cell receptor repertoires from bulk RNA-seq using TRUST4 to compare peripheral blood (n = 24), bronchoalveolar lavage (BAL; n = 55), and induced sputum (n = 100) from a well-characterized rural Ugandan cohort comprising individuals with and without HIV infection and chronic obstructive pulmonary disease (COPD). Comparative analyses revealed marked differences in productive TCR recovery, clonotype richness, diversity, and clonotype sharing across anatomical compartments, with induced sputum yielding fewer reconstructed clonotypes than BAL and peripheral blood despite comparable sequencing depth. In contrast, TRBV gene usage, CDR3 length distributions, amino acid composition, and global TRB sequence-space organization were broadly conserved. Together, these findings establish a comparative reference framework for transcriptome-derived TCR repertoire analysis across peripheral blood and the human respiratory tract.

Files

Steps to reproduce

TRUST4: Immune repertoire reconstruction from sputum bulk RNA-seq data. We reconstructed T cell receptor repertoires from bulk RNA-seq data using the TRUST4 algorithm28. TRUST4 processes data in three main stages: candidate read extraction, de novo assembly, and annotation. During candidate extraction, reads mapping to known TCR V, J, or constant (C) gene regions, as well as reads with significant k-mer overlap to these loci, were retained. Candidate reads were then assembled de novo into contigs using an overlap-based greedy extension strategy that prioritizes highly expressed receptor sequences. Assembled contigs were annotated by alignment to the IMGT reference database to assign V, J, and C genes and to identify complementarity-determining region 3 (CDR3) sequences. TRUST4 outputs were generated in AIRR-compliant CSV format for downstream analysis. TCR clonotype data sources and preprocessing. Bulk RNA-seq data from induced sputum, bronchoalveolar lavage (BAL), and peripheral blood samples were independently processed with TRUST4 to reconstruct TCR repertoires for each compartment. Sample-level clinical metadata, including HIV status, COPD status, and combined dual status, were imported from accompanying metadata files and merged with repertoire outputs. Raw TRUST4 outputs were filtered to retain only productive TCR rearrangements from the TRA, TRB, TRD, and TRG loci. Sequences lacking V gene, J gene, or junction amino acid (CDR3) annotations were excluded. Clonotypes were defined as unique combinations of V gene, J gene, and CDR3 amino acid sequence, and clonotype abundance was quantified using TRUST4-reported supporting read counts. Microbiome composition data were obtained from matched species- or genus-level abundance tables and aligned to TCR data at the sample level for integrative analyses. All downstream analyses were performed on filtered, AIRR-compliant clonotype tables with harmonized clinical and microbiome metadata. QUANTIFICATION TCR clonotype quantification. For each sample, T cell receptor (TCR) clonotypes were reconstructed from bulk RNA sequencing data using TRUST4 and quantified separately for the TRA, TRB, TRD, and TRG chains. Clonotype abundance was defined as the number of supporting reads per unique CDR3 amino acid sequence and V-J gene combination. Group-level summaries were generated by anatomical compartment (sputum, bronchoalveolar lavage [BAL], and peripheral blood) and further stratified by HIV status, COPD status, and combined dual status. Total and unique clonotypes were enumerated at both the sample and compartment levels. Repertoire diversity and richness. Within-sample repertoire diversity was quantified using Shannon entropy and Simpson diversity indices, calculated from clonotype frequency distributions. Clonal richness was defined as the number of unique clonotypes detected per sample.

Institutions

  • Makerere University College of Health Sciences
    Kampala, Kampala

Categories

RNA Sequencing, Amplicon Sequencing

Funders

Licence