Integrated Co-expression Analysis of Host–Parasite Transcriptomes Reveals Mechanisms of Host Modulation in an Ant–Cestode System

Published: 23 January 2026| Version 1 | DOI: 10.17632/85w27rg8m9.1
Contributors:
,
,

Description

How parasites interact with their hosts at the molecular level is a central question in parasitology, yet identifying host pathways directly targeted by parasites is challenging because infections often have broad effects on host physiology. This difficulty is particularly pronounced in non-model systems, such as the interaction between the parasitic tapeworm Anomotaenia brevis and its intermediate host, the ant Temnothorax nylanderi, in which infection induces strong phenotypic changes. Here, we integrated host and parasite transcriptomes through a combined weighted gene co-expression network analysis (WGCNA) to identify candidate genes and gene networks involved in this interaction. We detected strong negative correlations between parasite and host gene expression, whereas within-species associations were largely positive. Candidate parasite genes were associated with host molecular pathways relevant to infection and host phenotype. The gene networks and expression correlations identified were consistent with those described in model parasite–host systems, supporting the robustness of our approach. Besides, our analysis provided initial functional insights into previously unannotated parasite proteins that may act as effectors of host manipulation. Expression of these parasite genes was correlated with host genes involved in oxidative stress resistance, metabolism, muscle function, immunity, and cuticular sclerotization. These associations suggest that the parasite may modulate multiple host pathways to facilitate infection and transmission. Overall, our findings advance our understanding of molecular mechanisms underlying parasite interference and highlight the value of integrating host and parasite transcriptomic data. More generally, our combined WGCNA framework provides a useful tool for uncovering transcriptional interactions in complex host–parasite systems. Files and variables File: link_wgcna.R Description: R-script for the analysis of the transcriptome data File: merged_ant_cestode_gcm.csv Description: Gene count matrix of parasite and host transcriptomes combined, each column is a sample name, each row is a gene, if the gene name starts with a "g", it is a parasite gene, when it starts with an "M" it is a host gene. Code/software Software needed is R, R packages used are: WGCNA, ggplot2, dplyr, topGO, clusterProfiler, pathview, data.table, stringr, RColorBrewer, tidyverse, forcats Access information Data was derived from the following source: 10.5061/dryad.8cz8w9h3b Additional data can also be found on this link

Files

Steps to reproduce

We obtained our transcriptome data from Sistermans et al. (2025; methodological details see supplement [13]). Raw reads are published under the SRA database of NCBI (PRJNA1246159), and the gene count matrices, GO terms and KEGG terms on Dryad (DOI: 10.5061/dryad.8cz8w9h3b). We targeted the host’s fat body as in insects, this physiologically active organ is responsible for synthesizing and processing proteins essential for immunity, fecundity, and longevity [27]. Gene counts were paired per sample (table S1); as both originated from the same worker ant. This process yielded a joint gene count matrix for 15 samples, encompassing transcripts from both cestode and ant genes (Table S1). From this combined matrix, we filtered out genes with fewer than ten counts in at least five samples and verified the data for missing entries using the WGCNA package (version 1.72-2) [20] in R (version 4.3.2). To construct an unsigned co-expression network, we set the soft-thresholding power to 8. After testing various module sizes, we established a minimum module size of 100 and confirmed this using a TOM plot. Modules with a dissimilarity threshold of 0.2 were merged, and the result was validated with an additional TOM plot (Fig. S1). We identified hub genes by selecting the genes with the top 10% highest connectivity within their modules. We used a chi-square test per module to check whether there were more cestode hub genes than would be expected by chance. For this, we used the proportion of cestode genes per module as the expected variable, the proportion of cestode hub genes as the observed variable and the number of hub genes as the sample size. We corrected the p-values for multiple testing using Benjamini-Hochberg adjustment. We converted module-specific topological overlap matrices into Cytoscape objects (Cytoscape version 3.26.0) [28] to visualize weighted gene–gene correlations and construct parasite–host co-expression networks. These networks were used to test whether genes preferentially correlated with genes from the other species, relative to expectations under random association based on the proportion of ant and cestode genes within each module (e.g., a 50:50 ratio predicts equal proportions of correlations with genes from both species). Although gene expression within the same species is likely to be more strongly statistically linked, we tested against the more conservative null hypothesis of random correlations. We calculated the proportion of genes that are correlation partners of cestode genes for both ant and cestode genes and took their respective averages in each module, we then tested deviations from the null hypothesis using a chi-square test, with the number of ant or cestode genes in each module as the sample size. P-values were adjusted for multiple testing using the Benjamini-Hochberg method resulting in conservative estimates.

Institutions

Categories

Parasitology, Immune Response to Parasite, Weighted Gene Co-Expression Network Analysis, Molecular Parasitology

Funders

Licence