Viewport Size Code:
Login | Create New Account
picture

  MENU

About | Classical Genetics | Timelines | What's New | What's Hot

About | Classical Genetics | Timelines | What's New | What's Hot

icon

Bibliography Options Menu

icon
QUERY RUN:
HITS:
PAGE OPTIONS:
Hide Abstracts   |   Hide Additional Links
NOTE:
Long bibliographies are displayed in blocks of 100 citations at a time. At the end of each block there is an option to load the next block.

Bibliography on: Pangenome

The Electronic Scholarly Publishing Project: Providing world-wide, free access to classic scientific papers and other scholarly materials, since 1993.

More About:  ESP | OUR CONTENT | THIS WEBSITE | WHAT'S NEW | WHAT'S HOT

ESP: PubMed Auto Bibliography 07 Aug 2026 at 01:34 Created: 

Pangenome

Although the enforced stability of genomic content is ubiquitous among MCEs, the opposite is proving to be the case among prokaryotes, which exhibit remarkable and adaptive plasticity of genomic content. Early bacterial whole-genome sequencing efforts discovered that whenever a particular "species" was re-sequenced, new genes were found that had not been detected earlier — entirely new genes, not merely new alleles. This led to the concepts of the bacterial core-genome, the set of genes found in all members of a particular "species", and the flex-genome, the set of genes found in some, but not all members of the "species". Together these make up the species' pan-genome.

Created with PubMed® Query: ( pangenome[TIAB] OR "pan-genome"[TIAB] OR "pan genome"[TIAB] ) NOT pmcbook NOT ispreviousversion

Citations The Papers (from PubMed®)

-->

RevDate: 2026-08-05
CmpDate: 2026-08-05

Ortiz-Flores C, Romero-Rodríguez A, Villanueva-Enríquez R, et al (2026)

Comparative genomic analysis of Clostridioides difficile strains in Mexico: insights into virulence and resistance.

Access microbiology, 8(7):.

Clostridioides difficile infection (CDI) remains a major global health threat due to the emergence of hypervirulent, multidrug-resistant lineages. However, the evolutionary dynamics and resistance-associated genomic profiles of strains circulating in underrepresented regions, such as Mexico, remain poorly characterized. Here, we present a comprehensive genomic and phylogenetic analysis of 77 Mexican C. difficile strains compared with 74 strains from other parts of the world. Using whole-genome sequencing and core-genome MLST, we identified 19 sequence types (STs) grouped across 3 clades, with hypervirulent ST01 dominating clade 2. Virulome analysis showed conserved toxin gene profiles (tcdA, tcdB and cdtAB) across strains, while clade-specific differences were observed in adhesion and survival genes. These variations, particularly pronounced in clade 2 strains from both global and Mexican collections, may contribute to enhanced persistence and transmissibility. Pangenome analysis of 151 genomes highlighted distinct genomic architectures. Clade 2, enriched in ST01 epidemic lineages, contained 5,480 genes (58% core, 42% accessory), showing a compact structure consistent with recent clonal expansion. In contrast, clade 1 displayed the highest diversity, with 8,584 genes (30% core, 70% accessory), indicative of an open and dynamic pangenome, while clade 4 showed a smaller, more conserved profile (4,784 genes, 63% core). These findings underscore the contrasting evolutionary strategies among clades. Notably, Mexican ST01 strains exhibited a distinct resistome, including the high prevalence of the vanG operon and the VanR T115A substitution (94% vs. 23% globally), as well as near-complete prevalence of the PnimB[G] mutation associated with reduced metronidazole susceptibility. This pattern may reflect local selective pressures associated with antimicrobial exposure. Phenotypic susceptibility testing of newly sequenced isolates showed that most ST01 strains remained susceptible to metronidazole and vancomycin despite carrying resistance-associated determinants. Our findings highlight the urgent need to recognize hypervirulent and resistant C. difficile lineages arising outside traditional surveillance regions. These Mexican strains not only reflect regional antibiotic usage patterns but also represent a potential reservoir of globally significant resistance traits. This work underscores the importance of integrating genomic surveillance across all continents to refine treatment protocols, prevent outbreaks and contain the spread of resistant CDI.

RevDate: 2026-08-05
CmpDate: 2026-08-05

Wen H, Bai Y, Guo Z, et al (2026)

Pangenome and Pan-Transcriptome Analysis of the WOX Gene Family Reveals Evolutionary Conservation and Diversification in Alfalfa.

Ecology and evolution, 16(8):e74048.

The WUSCHEL-related homeobox (WOX) gene family encodes plant-specific transcription factors that play pivotal roles in meristem maintenance, organogenesis, regeneration, and developmental phase transitions. Despite their importance, the pangenome-scale composition, expansion dynamics, and evolutionary constraints of WOX genes in alfalfa (Medicago sativa) remain poorly understood. Here, we performed an integrated pangenome and pan-transcriptome analysis of the WOX gene family using 24 high-quality alfalfa genomes. A total of 432 WOX genes were identified and clustered into 24 orthologous gene groups (OGGs). Pangenome profiling revealed that the WOX family is highly conserved across alfalfa accessions, with members of the Ancient clade showing greater conservation than those of the Intermediate and WUS clades. Duplication pattern analysis indicated that dispersed duplication was the primary force driving WOX family expansion, whereas whole-genome duplication (WGD) events predominantly contributed to the long-term retention of core and softcore members. Pairwise Ka/Ks analysis indicated strong purifying selection on most WOX gene pairs, whereas wild-germplasm genes and Intermediate-clade members showed elevated selective pressures. Notably, MsWOX9-related homologs had exceptionally high Ka/Ks values, suggesting potential functional divergence linked to agronomically relevant traits. Furthermore, pan-transcriptome analysis revealed that WOX genes could be classified into three distinct expression patterns during leaf senescence. Collectively, this study presents the first comprehensive pangenome- and pan-transcriptome-based characterization of the WOX family in alfalfa, providing new insights into its evolutionary dynamics and expression divergence.

RevDate: 2026-08-05

Sun N, Chen Y, Wu X, et al (2026)

Comparative genomic characterization and antimicrobial resistance of bacteremia-causing Enterococcus faecium and Enterococcus faecalis in a Chinese hospital.

Microbiology spectrum [Epub ahead of print].

Enterococci are common commensals of the human gut and important opportunistic pathogens, with Enterococcus faecium and Enterococcus faecalis being the most clinically prevalent species. A significant epidemiological shift has emerged with an increasing clinical burden of E. faecium. To compare genomic evolution of E. faecium and E. faecalis, we performed whole-genome sequencing on 93 E. faecium and 32 E. faecalis isolates causing bloodstream infections at a single hospital (2022-2024). Analysis of patient demographics revealed that E. faecium infections originated from fewer sources than E. faecalis, with a higher proportion deriving from intra-abdominal infections. Multilocus sequence typing identified ST78 and ST789 as the predominant sequence types for E. faecium, whereas ST16 and ST179 were most common for E. faecalis. E. faecium carried more antimicrobial resistance genes and putative virulence marker (PVM)-type virulence genes than E. faecalis, with vancomycin resistance predominantly mediated by vanHAX (33/93, 35.5%) and a single E. faecalis isolate also carrying vanHAX (1/32, 3.1%); the structurally incomplete vanHMX gene cluster was detected in 11 E. faecium isolates. Pan-genome analysis indicated a larger core genome in E. faecalis compared to E. faecium, consistent with greater plasmid replicon diversity in the latter. Intra-host comparisons showed that two E. faecalis pairs from the same patient were clonally related, with one isolate acquiring a vanHAX plasmid conferring vancomycin resistance. In contrast, E. faecium isolates exhibited marked genomic diversity even among clonally related pairs. These findings suggest that E. faecium possesses greater genomic plasticity and adaptive potential to the clinical environment.IMPORTANCEThis study provides a detailed comparison of clinical and genomic features between Enterococcus faecium and Enterococcus faecalis from the same hospital setting. We show that E. faecium isolates, mainly ST78/ST789, carry more antimicrobial resistance genes and a higher number of putative virulence marker (PVM) genes than E. faecalis, reflecting their hospital-adapted nature. E. faecium also exhibits a smaller core genome and greater diversity of plasmid replicon types, indicating higher genomic plasticity and capacity for horizontal gene transfer. By contrast, E. faecalis retains a larger core genome and a set of classical virulence factors, and its within-host isolates are clonally related. These distinct genomic profiles help to understand how the two species adapt to clinical environments and may inform more targeted infection control strategies and resistance surveillance.

RevDate: 2026-08-06
CmpDate: 2026-08-06

Zhang D, Ding Y, Kang W, et al (2026)

Comprehensive genomic characterization of extraintestinal pathogenic Escherichia coli isolated from neonates: multiple center insights into virulence, resistance, and transmission dynamics.

Genome medicine, 18(1):.

BACKGROUND: Neonatal extraintestinal pathogenic Escherichia coli (ExPEC), which can cause severe long-term sequelae by systemic infections, is gradually becoming the primary pathogen threatening neonatal health. The lack of large-scale genomic epidemiological investigation hinders further understanding of neonatal ExPEC. We conducted this nationwide multicenter study to support further strategies for improving neonatal ExPEC management.

METHODS: The neonatal ExPEC strains and clinical information, including antimicrobial resistance phenotype, were collected from nine centers within 7 provinces across China between 2018 and 2023. Whole-genome sequencing was performed. Sequence types (ST) and serotypes were acquired to characterize the strains. Phylogenetic analysis and pan-genomic analysis were conducted to identify the population structure. Bioinformatics analysis associated with virulence factors, antimicrobial resistance genes, and mobile genetic elements were conducted. To characterize the situation of horizontal gene transfer, we developed a computational tool for identifying horizontal evolutionary patterns from large-scale genomic draft assemblies. Co-occurrence and co-localization metrics were used to describe the synergistic effects and transmission mechanism of genes.

RESULTS: A total of 411 neonatal ExPEC strains were included. ST1193 (18·0%) was the main ST, while O75 (15·8%) was the most common serotype. Virulence factors and antimicrobial resistance genes were widely distributed across various STs, provinces, years, and isolation sites. Co-occurrence analysis revealed multiple clusters of virulence factors and antimicrobial resistance genes, suggesting co-transmission or co-evolution. Multiple kinds of mobile genetic elements were widely distributed throughout the country. The predicted plasmid-derived contig, genome islands, prophages, and transposons carry different pathogenic genes, respectively. Multiple pathogenic genes exhibited co-occurrence with a specific plasmid replicon, suggesting the critical role of plasmids in the evolution of ExPEC.

CONCLUSIONS: Our findings indicate that neonatal ExPEC had a shared phylogenetic spectrum with adult ExPEC isolates, but distinct dominant subtypes. Multiple virulence factors and drug resistance genes form a complex network that enhances pathogenicity. The formation of these gene clusters is associated with both the inherent genetic factors of ExPEC and the involvement of complex mobile genetic elements. These data accelerate the understanding of neonatal ExPEC, revealing the distribution of STs, serotypes, pathogenic genes, and transmission dynamics.

RevDate: 2026-08-06
CmpDate: 2026-08-06

Xie J, Liu T, Zhang M, et al (2026)

Cutibacterium nasicola sp. nov., isolated from nasal swabs of coal miners.

International journal of systematic and evolutionary microbiology, 76(8):.

Two Gram-stain-positive, catalase-positive, oxidase-negative, non-spore-forming, aerotolerant anaerobic and non-motile short rod-shaped bacterial strains, designated V947[T] and V970, were isolated from nasal swab samples of coal miners. Phylogenetic analyses based on 16S rRNA gene sequences and whole-genome sequences revealed that strains V947[T] and V970 represent a distinct lineage within the genus Cutibacterium, most closely related to Cutibacterium avidum ATCC 25577[T] (16S rRNA gene sequence similarity of 97.09%). Whole-genome comparative analyses showed that the average nucleotide identity values between the two strains and all validly published species of the genus Cutibacterium with correct nomenclature ranged from 76.55 to 90.26%, while the digital DNA-DNA hybridization values ranged from 21.80 to 40.40%, both of which are well below the accepted thresholds for species delineation. Pangenome analysis identified 711 species-specific gene clusters present in strains V947[T] and V970 but absent from all reference type strains of the genus, which were predominantly involved in carbohydrate transport and metabolism, inorganic ion transport and signal transduction mechanisms, suggesting distinct metabolic capabilities and potential for environmental adaptation. These genomic features not only support the delineation of a novel species but also expand the known genomic and functional diversity within the genus Cutibacterium. The predominant cellular fatty acids were iso-C15:0 and anteiso-C15:0. The major polar lipids of strain V947[T] were diphosphatidylglycerol, phosphatidylinositol and phosphatidylcholine, and the predominant menaquinones were MK-8(H4) and MK-9(H6). On the basis of phylogenetic, genomic, chemotaxonomic and phenotypic characteristics, strains V947[T] and V970 represent a novel species of the genus Cutibacterium, for which the name Cutibacterium nasicola sp. nov. is proposed. The type strain is V947[T] (=CGMCC 1.5844[T]=KCTC 59608[T]).

RevDate: 2026-08-04
CmpDate: 2026-08-04

Qi W, Wang J, Chang J, et al (2026)

Pan-genome analysis reveals structural variation- associated expression and evolutionary diversity of the ZmCYP450 gene family in maize.

Frontiers in plant science, 17:1874015.

The plant cytochrome P450 (CYP450) superfamily plays a key role in metabolic diversity and environmental adaptation; however, systematic analyses of its intraspecific structural variation, copy number dynamics, and evolutionary mechanisms remain limited. Using 27 high-quality maize reference genomes, we performed a pan-genomic analysis of the ZmCYP450 family, identifying 282 orthogroups (OGs) and 7,282 genes. The family exhibits a pattern of predominantly conserved genes with localized expansions and open pan-genome properties. PAV and CNV analyses revealed extensive gene deletions and copy number fluctuations outside core OGs, reflecting substantial intraspecific structural diversity. Analysis of duplication types and local colinearity suggested that proximal and dispersed duplications are the primary contributors to drive family expansion. Ka/Ks analysis indicated that most OGs are under purifying selection, while a subset shows evidence of positive selection. Further integration of structural variation and transcriptomic data suggested that SVs may affect gene function through mechanisms such as gene deletion, protein truncation, and remodeling of regulatory elements, suggesting a dual role of potential loss of function and expression modulation. Despite a relatively stable overall copy number, structural variant categories ('Typical', 'Atypical', and 'Missing') are widespread, and expression levels do not always correlate with copy number, suggesting a complex regulatory patterns. Tissue-specificity analysis revealed a large number of highly specific ZmCYP450 genes involved in secondary metabolism and environmental responses, with distinct expression profiles across different genetic backgrounds. This study provides a comprehensive pan-genomic view of structural variation, evolutionary patterns, and expression regulation in the ZmCYP450 family, offering a foundation for future functional studies and maize trait improvement.

RevDate: 2026-08-04

Geng Y, Abideen MZU, Shaukat A, et al (2026)

SNP discovery and applications in plant genetics for sustainable food security: A review.

International journal of biological macromolecules pii:S0141-8130(26)03858-4 [Epub ahead of print].

Single nucleotide polymorphisms represent the most abundant form of genetic variation in plant genomes and have become fundamental markers in modern plant genetics and breeding. Increasing global demand for food, combined with the pressures of climate change, environmental stress, and declining arable land, requires accelerated crop improvement strategies that exceed the capacity of conventional breeding approaches. In this context, SNP-based genomic technologies provide high-resolution tools for analyzing genetic diversity, identifying trait-associated loci, and enhancing selection efficiency in breeding programs. This review provides a comprehensive synthesis of SNP discovery methodologies, tracing their development from early Sanger sequencing approaches to advanced next-generation sequencing technologies, including whole-genome resequencing, genotyping-by-sequencing, and high-density SNP arrays. The article further examines the diverse applications of SNP markers in plant genetics, including genetic diversity analysis, linkage mapping, genome-wide association studies, marker-assisted selection, genomic selection, and evolutionary research. Key analytical and technical challenges, particularly those related to polyploid genome complexity, large-scale genomic data processing, and accurate variant interpretation, are critically discussed. In addition, emerging developments such as graph-based pangenomes, long-read sequencing technologies, machine learning-assisted SNP prioritization, and multi-omics integration are highlighted as promising directions for future research. By integrating recent technological advances with established genomic approaches, this review emphasizes the central role of SNP-based genomics in accelerating crop improvement and enabling climate-resilient, sustainable agricultural systems that support global food security.

RevDate: 2026-08-04
CmpDate: 2026-08-04

Jiang S, Li S, Han S, et al (2026)

Near telomere-to-telomere genome assembly of the rainbow trout (Oncorhynchus mykiss).

Scientific data, 13(1):.

The rainbow trout (Oncorhynchus mykiss) exhibits extensive karyotypic diversity (2n = 58-64) driven by Robertsonian translocations, yet widely used reference genomes are derived from North American lineages, leaving Chinese aquaculture populations underrepresented. Here, we present a near telomere-to-telomere (T2T) genome assembly of a farmed rainbow trout from China. Integrating PacBio HiFi, ONT ultra-long reads, and Hi-C data, we assembled a 2.29 Gb genome with 99.04% anchored to 30 chromosomes. Notably, the genome contains only 18 gaps, with 16 gap-free chromosomes and 13 achieving T2T status. Comparative synteny analysis revealed a third chromosomal fission/fusion iteration in which Swanson Omy14 splits into Arlee Omy14 and Omy32. Annotation identified 43,137 protein-coding genes, with a BUSCO completeness of 98.9%. This dataset provides a valuable resource for resolving lineage-specific structural variation, supporting pangenome construction and facilitating molecular breeding in rainbow trout.

RevDate: 2026-08-05

Lee K, Korani W, Pokhrel S, et al (2026)

Long-read low-pass sequencing enhances variant detection in a peanut MAGIC population.

G3 (Bethesda, Md.) pii:8751453 [Epub ahead of print].

Accurate genotyping accelerates crop improvement, yet long-read sequencing remains underused in breeding due to cost. We present a scalable long-read low-pass (LRLP) sequencing framework for high-throughput variant discovery and trait mapping. Using PacBio HiFi reads in an allotetraploid peanut (Arachis hypogaea; AABB, 2n = 4x = 40) MAGIC population, we generated both LRLP and short-read low-pass (SRLP) data. At comparable depths, LRLP achieved substantially greater whole-genome and gene-space coverage than SRLP. Data were analyzed using both a single-reference genome and an 18-parent pangenome graph constructed with KhufuPan, a new tool for graph-based genotyping. Across analytical approaches, LRLP consistently identified more SNPs, indels (2-1,000 bp), and structural variants (>1 kb) than SRLP, improving genotype resolution and selection accuracy, particularly for large structural variants. By reducing cost barriers and increasing variant discovery in complex genomes, LRLP provides a practical path for deploying advanced genomics in under-resourced and orphan crops critical to global food security.

RevDate: 2026-08-05
CmpDate: 2026-08-05

Thanh LU, Tan VN, Huyen TTN, et al (2026)

Genomic characterization of Bacillus velezensis BP5 and BP103: insights into biocontrol potential against bacterial leaf spot on pepper.

Frontiers in microbiology, 17:1757903.

Bacillus velezensis strains BP5 and BP103, isolated from pepper rhizosphere soil in the Mekong Delta, Vietnam, displayed potent antagonistic activity against Xanthomonas euvesicatoria, which causes pepper bacterial spot. In dual-culture assays, BP5 formed the largest inhibition zone (36.7 mm), outperforming BP103 (32.1 mm) and oxolinic acid (31.3 mm). Both Gram-positive, rod-shaped strains produced extracellular protease, lipase, amylase, cellulase, and siderophores. Whole-genome sequencing revealed compact and genetically stable genomes: BP5 (4.08 Mb, 46.04% GC, 4,047 protein-coding genes, 5 prophages) and BP103 (3.91 Mb, 46.46% GC, 3,793 genes, 2 prophages). Phylogenetic reconstruction and average nucleotide identity (ANI of more than 98%) confirmed their assignment to B. velezensis, with closest relatedness to strain 160. Orthologous gene analysis across 28 B. velezensis genomes and B. subtilis DSM10 identified 2,593 single-copy genes. Compared to BP103, BP5 harbors greater numbers of multi-copy orthologs, other orthologs, and unique paralogs (388, 910, and 5 in BP5 with 336, 763, and 1 in BP103, respectively). Pangenome analysis showed that the core genome size decreased progressively with the addition of new genomes; BP5 showed more pronounced genetic divergence driven by expanded unique paralogs and shell gene families (with 456 in BP5 compared with 203 in BP103). Both strains encode 13 secondary metabolite biosynthetic gene clusters, including surfactin, fengycin, bacillaene, macrolactin, difficidin, bacilysin, bacillibactin, mersacidin, and locillomycin, plus four previously undescribed terpene/PKS. The CAZyme repertoire was richest in BP5 (133 genes), exceeding BP103 and the commercial strain FZB42 (both 129 genes), thereby enhancing adhesion and adhesion and rhizosphere colonization. These findings establish BP5 and BP103 as highly promising biocontrol agents, combining high genetic stability with diverse secondary metabolite profiles, suitable for development into sustainable microbial bio-bactericides.

RevDate: 2026-08-01
CmpDate: 2026-08-01

Lucas JK, Hebbar P, Liao WW, et al (2026)

HPRC2: A human pangenome reference with near-complete coverage of common genetic variation.

bioRxiv : the preprint server for biology pii:2026.07.21.739710.

A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.

RevDate: 2026-08-01
CmpDate: 2026-08-01

Cohen ZP, Perkin L, Frandsen PB, et al (2026)

Bacterial DNA invasion triggers transposable element proliferation and genome expansion.

bioRxiv : the preprint server for biology pii:2026.07.14.738529.

Genome size variation in eukaryotes is driven largely by transposable elements (TEs), yet the biological mechanisms that initiate their proliferation remain understudied. Here, we identify a recurrent association between bacterial horizontal gene transfer (HGT) and bursts of TE activity that contribute to genome expansion. By leveraging comparative genomics and genus-level pangenome analyses across three species of the nut weevil, Curculio , we detect extensive bacterially derived DNA sequences embedded within structurally dynamic genomic regions. These HGT-associated regions are dominated by a small number of young, proliferating TE families, particularly DNA type II Mavericks, which encapsulate transferred bacterial sequences and comprise a substantial fraction of recent genomic DNA in derived lineages. Analyses of codon usage bias, intron length, and functional enrichment suggest that most transferred genes undergo progressive pseudogenization over evolutionary time, whereas a subset of selectively advantageous HGTs persist. Together, our findings support a model linking foreign DNA invasion with TE proliferation, genome size variation, and molecular innovation.

RevDate: 2026-08-01

Rosenbaum SW, Escalona M, May SA, et al (2026)

A Mid-Atlantic Brook Trout (Salvelinus fontinalis) genome to advance conservation and comparative genomics.

G3 (Bethesda, Md.) pii:8748888 [Epub ahead of print].

Brook Trout (Salvelinus fontinalis) are experiencing genomic erosion and demographic declines across the southern portion of their native distribution. A regionally representative reference genome is necessary to support conservation genomic initiatives for this species. While Brook Trout exhibit substantial phylogenetic structure across their range, only a single reference genome from Northeastern Canada (ASM2944872v1) is currently used to guide range-wide analyses. Consequently, reference bias is expected to cause erroneous sequence alignment and spurious variant detection for Brook Trout from divergent lineages. To prevent reference bias and invalid inference when studying Mid-Atlantic Brook Trout, we assembled a de novo, chromosome-level genome using a wild individual from Virginia, USA and produced a gene annotation with publicly available RNA-seq data. The assembly combined three complementary sequencing types (PacBio HiFi, Oxford Nanopore Ultra-Long, and Dovetail Omni-C) and benefited from manual curation using the Pretext software suite. The 2.91 Gb genome, known as mSalvFont1.0, consists of 3,333 contigs (N50 = 3.3 Mb) and 1,299 scaffolds (N50 = 55 Mb), with 81% of the assembly contained within 42 chromosome-scale scaffolds. Indices of accuracy and completeness reveal a quality value of 59 (99.999% accuracy), k-mer completeness of 95%, and 99.3% complete BUSCOs from the Actinopterygii orthologous gene set. Additionally, we identified reference bias (26.94% heterozygosity inflation) when variant detection of a Mid-Atlantic-origin sample relied on alignment to ASM2944872v1, instead of mSalvFont1.0. We anticipate that mSalvFont1.0 will increase the accuracy and precision of regional conservation genomic initiatives and expedite the genesis of a Brook Trout pangenome.

RevDate: 2026-08-01
CmpDate: 2026-08-01

Shen J, Fu Q, Macias-Velasco JF, et al (2026)

Pansoma, a machine learning tool for identifying somatic variants using pangenome graphs.

bioRxiv : the preprint server for biology pii:2026.05.27.726245.

Somatic variant calling, the identification of mutations in non-germline cells acquired over an individual's lifetime, is critical for studying diseases, including cancer, and for developing precision oncology strategies. Traditional somatic variant calling methods rely on linear reference genomes, which do not adequately capture human genetic diversity and result in reference bias, compromising the accuracy of somatic variant detection. Recently developed graph-based human pangenome reference represents diverse genetic variants across human populations and has promised to drive advances in many genetics and genomics studies. In this study, we introduced Pansoma, a novel pangenome-native and machine learning-based tool specifically designed for somatic variant calling using a pangenome graph reference. Pansoma performs somatic variant detection from both short- and long-read sequencing data by learning tensor representations of alignment on graph nodes rather than on a linear reference. Pansoma outputs variant representations anchored to the pangenome graph paths and conventional somatic variant calls remapped to the linear reference. Additionally, we provide accompanying bioinformatics tools tailored for graph-based genomic data management and variant calling results analysis. Benchmarking shows that Pansoma not only improves tumor-only somatic variant detection but also preserves graph-specific variant representations that are not directly recoverable from linear- reference outputs.

RevDate: 2026-07-30
CmpDate: 2026-07-30

Meng F, Li R, Yu Y, et al (2026)

Cross-host transmission of Riemerella anatipestifer to chickens: Genomic evolution and identification of the novel vapX-like-vapD toxin-antitoxin system.

Virulence, 17(1):2711521.

Riemerella anatipestifer (R. anatipestifer), a well-known waterfowl pathogen, increasingly threatens Chinese poultry by spreading to chickens. The genetic differentiation and host adaptation following cross-host transmission remain unclear. Here, we characterized a highly virulent, multidrug-resistant chicken-source strain (SDAU-RA1) and performed comparative genomics with global R. anatipestifer strains to elucidate population structure and evolutionary dynamics. SNP phylogeny revealed significant geographic clustering and dominant clonal groups. Strains from different hosts showed a pattern of "overall mixing with local clustering," and ancestral state reconstruction (ASR) identified multiple independent duck-to-chicken spillover events, confirming cross-host transmission rather than strict host-specific evolution. In terms of virulence, certain virulence genes are enriched specifically in chicken-source strains. Notably, the study is the first to identify and confirm vapX-like-vapD as a functional type II toxin-antitoxin system in chicken-source R. anatipestifer, demonstrating that it enhances biofilm formation, intracellular survival, and antibiotic persistence. Analysis of the geographical distribution and temporal dynamics of antibiotic resistance genes (ARGs) reveals high heterogeneity among R. anatipestifer strains from different hosts. Pangenome analysis revealed that R. anatipestifer possesses an open pangenome, conferring high genetic plasticity. In conclusion, our study shows R. anatipestifer transmits to chickens without strict host-specific adaptation, though incipient genetic differentiation has emerged. The discovery of plasmid pRASD and its carried vapX-like-vapD system suggests key mechanisms for the adaptive evolution and enhanced pathogenicity of R. anatipestifer. These findings enhance our understanding of cross-host transmission and underscore the importance of continuous surveillance of chicken-source R. anatipestifer and its novel mobile genetic elements.

RevDate: 2026-07-30
CmpDate: 2026-07-30

Xiao C, Jiang S, Guo Z, et al (2026)

Integrating genomic and transcriptomic analyses reveals virulence factors and bloodstream infection mechanisms of Acinetobacter junii.

Microbial genomics, 12(7):.

Acinetobacter junii, an emerging pathogen of the Acinetobacter genus, exhibits concerning environmental persistence and clinical virulence potential. Despite its growing significance in bloodstream infections (BSIs), the genetic determinants distinguishing pathogenic from environmental strains remain poorly characterized, limiting our understanding of its transmission dynamics and host adaptation mechanisms. We explored the diversity and population structure of A. junii using a comprehensive genome dataset, including four strains isolated from China, one of which (acj0210) was derived from BSIs. We analysed the pangenome of A. junii and characterized its virulence and antibiotic resistance profile. Bacterial genome-wide association studies (bGWAS) were performed to identify the virulence potential. The virulence potential for BSIs was further investigated through RNA-seq analysis under different conditions. Additionally, transmission dynamics of bla NDM-carrying plasmids in carbapenem-resistant A. junii were inferred from plasmid network analysis. We identified seven lineages within A. junii, with BAPS1 as the dominant group. Discrepancies in accessory genes were observed between lineages, along with unique genes, particularly one novel capsular polysaccharide variant in acj0210. Using bGWAS, we uncovered key genetic differences, including transporter-encoding genes (tauA, urtD), regulatory mutations and exopolysaccharide-encoding genes that distinguish human and environmental strains. Transcriptomic analysis revealed potential virulence factors in BSIs, with distinct gene expression patterns between blood and immune cells, despite shared fatty acid metabolism profiles. As critical resistance determinants, bla NDM genes may have been acquired from A. baumannii. A. junii has adapted to diverse geographical and ecological niches, particularly in BSIs, with lineage-specific virulence factors aligned with antibiotic resistance genes, especially those carried by bla NDM harbouring plasmids.

RevDate: 2026-07-31
CmpDate: 2026-07-31

Balogun IO, Mancuso CP, TD Lieberman (2026)

High-precision binary trait association on phylogenetic trees.

Microbial genomics, 12(7):.

Traditional methods for identifying associations between genomic features and traits, or between pairs of genomic traits, struggle when applied to bacterial genomes. While several microbial genome-wide association study (mGWAS) methods have been developed to account for the fact that genome-wide linkage in bacteria creates strong evolutionary-induced associations, these methods have high false discovery rates or lack statistical power, have poor performance on negative interactions and face computational limits at the scale required for pangenome-wide study of gene-gene interactions. Here, we present Simulation-based Phylogenetic iNteraction Inference (SimPhyNI), a computationally optimized framework for efficient and rigorous mGWAS studies. SimPhyNI builds null co-occurrence distributions by independently simulating traits using phylogenetically informed parameters, novelly including time to first event. The constrained variation in these simulations, combined with log odds ratio scoring for comparing across traits, robustly identifies both positive and negative associations. Using synthetic datasets mimicking both gene-gene and gene-trait associations, we demonstrate that SimPhyNI achieves high precision and recall for both positive and negative interactions. We demonstrate SimPhyNI's utility by detecting interactions between phage defence systems in Escherichia coli and gene-gene interactions across the entire E. coli pangenome (>9 million tests). Though developed here for binary traits, SimPhyNI's design supports extension to multi-state and continuous traits using generalized models of stochastic simulation. SimPhyNI's performance and scalability enable genome-wide discovery of genetic interactions that drive microbial function, ecology and disease.

RevDate: 2026-08-01
CmpDate: 2026-08-01

Megalovasilis G, Bochalis E, Chartoumpekis D, et al (2026)

Z-DNA-induced genomic instability in the human pangenome.

Research square pii:rs.3.rs-10129918.

Z-DNA is a non-canonical, left-handed nucleic acid conformation with roles in gene regulation, genomic variation, and genome instability. Recent advances in long-read sequencing enable systematic interrogation of Z-DNA in repetitive and previously unresolved parts of the human genome. Leveraging the complete telomere-to-telomere reference human genome and 464 haplotype-resolved genome assemblies from individuals of multiple ethnicities, we systematically mapped Z-DNA-forming sequences and investigated their genomic and mutational landscape. We show that Z-DNA density is highly constrained across human haplotypes and superpopulations, with substantial enrichment in repetitive regions absent or unresolved in GRCh38, including centromeric, pericentromeric, and acrocentric loci. By analyzing more than 52 million variants across multiple mutation categories in the human pangenome, we show that Z-DNA loci are strongly enriched for mutational events not explained by sequence composition alone. The strongest associations were observed for small insertions, deletions, and complex structural variants. Finally, we replicated our findings using more than 250,000 de novo mutations. Together, these findings establish Z-DNA as a functional element of the human genome, with important implications for genome instability and human disease.

RevDate: 2026-07-29
CmpDate: 2026-07-29

Nguyen CTK, Dinh HT, Dang DQB, et al (2026)

Genomic and pan-genomic analyses of Bacillus subtilis B13 provide insights into biosynthetic potential and genetic traits associated with environmental adaptation.

Antonie van Leeuwenhoek, 119(8):.

This study describes the genomic features of Bacillus subtilis B13 (= VTCC 910231) to elucidate the genetic basis underlying its reported antimicrobial activity, metabolic versatility, and environmental adaptability. The draft genome comprised 4,349,051 bp with a GC content of 43.5% and 4436 predicted coding sequences. Genome-based analyses assigned B13 to B. subtilis subsp. subtilis, supported by high average nucleotide identity (98.3%) and digital DNA-DNA hybridization (99.7-99.8%) values. Genome mining identified one gene cluster encoding an unidentified sactipeptide along with seven biosynthetic clusters involved in the production of compounds with potential antibacterial activity, including fengycin, surfactin, bacillaene, bacillibactin, bacilysin, subtilosin A, and sporulation-killing factor. These clusters may contribute to its observed bioactive properties. Comparative pan-genome analysis suggested an open genomic architecture dominated by accessory genes, with B13 harboring 67 unique gene clusters at the species level and 336 strain-specific gene clusters in a niche-focused dataset, most of which remain functionally uncharacterised. The annotated genes are associated with environmental adaptation. The genome revealed mobile elements, indicating genome plasticity and potential horizontal gene transfer, but no plasmids were detected. Three high-confidence genomic islands (251 kb, 5.8% of the genome) contained mobility-related genes but lacked a virulence gene cluster and antibiotic resistance genes. Functional profiling explored a collection of genes associated with stress response, signal transduction, transport, motility, chemotaxis, and DNA repair. These findings provide insights into genomic features related to the biosynthetic potential, genomic plasticity, and safety profile of B13, and suggest putative determinants of environmental adaptation, while reflecting pan-genome diversity in strain-specific traits.

RevDate: 2026-07-29

Dos Reis JBA (2026)

Functional diversity and ecological consequences of endophytic Bacillus-plant interactions.

Folia microbiologica [Epub ahead of print].

The genus Bacillus, particularly endophytic species, has been widely studied as a source of plant growth-promoting bacteria in agricultural systems. These microorganisms contribute to plant performance through nutrient acquisition, phytohormone production, pathogen suppression, microbiome modulation, and enhanced tolerance to biotic and abiotic stresses. However, their ecological roles, functional plasticity, and genomic diversity remain poorly integrated into conceptual frameworks that extend beyond crop-based applications. Functional plasticity is reflected in their ability to colonize diverse plant hosts and tissues and to promote similar plant responses through distinct molecular mechanisms. Likewise, genomic diversity is evidenced by variation in accessory genomes, biosynthetic gene clusters, and regulatory networks that shape ecological functions and metabolite production. This review examines endophytic Bacillus as a model for understanding how metabolically versatile and genomically plastic bacteria establish functional, but context-dependent, associations with plants. Drawing on evidence from functional genomics, pangenomics, metabolomics, and microbial ecology, we discuss mechanisms associated with plant growth promotion and emphasize their dependence on host identity, environmental conditions, and microbial interactions. We address functional convergence arising from distinct genetic and metabolic routes, the contribution of accessory genomes and regulatory variation, and the ecological consequences of microbial inoculation in resident plant-associated microbiomes. We also highlight the limitations of in vitro screening approaches and the need for experimental validation across multiple biological scales to establish robust genotype-phenotype relationships. Finally, we extend the discussion beyond agricultural systems to consider the use of endophytic Bacillus in wild plant systems and ecological restoration, emphasizing the importance of evaluating both functional outcomes and ecological impacts.

RevDate: 2026-07-29

Gao S, Oshima KK, Chuang SC, et al (2026)

A global view of human centromere variation and evolution.

Nature [Epub ahead of print].

Centromeres are essential chromosomal regions that ensure accurate chromosome segregation during cell division, yet their highly repetitive sequence has historically hindered their complete assembly and characterization[1]. Consequently, the full spectrum of centromere diversity across individuals, populations and evolutionary contexts remains largely unexplored. Here we address this gap in knowledge by assembling and characterizing 2,110 centromeres from diverse individuals representing 5 continental and 28 population groups. Using bioinformatic tools tailored for centromeres, we identify variation, including 226 centromere haplotypes and 1,870 α-satellite higher-order repeat variants. While most centromeres have a single kinetochore site, we find that around 6% have di-kinetochores, and less than 1% have tri-kinetochores, which we confirm using long-read chromatin profiling and multigenerational inheritance. We also show that kinetochore position is closely associated with the underlying sequence and structure of the centromere. To understand the nature of evolutionary change, we compared these centromeres to 5,747 centromeres assembled by the Human Pangenome Reference Consortium. We show that centromeres have a 20-fold variation in mutation rate, and a subset of centromeres has evidence of archaic hominin introgression. We validate these mutation rates in a 4-generation, 28-member family and show that the kinetochore site is the most rapidly mutating region in the centromere. We propose a model that reveals an 'arms race' between centromeric sequence and proteins, with frequent mutations within the kinetochore site that lead to changes in genetic and epigenetic landscapes and, ultimately, rapid evolution of these critically important regions.

RevDate: 2026-07-29
CmpDate: 2026-07-29

Marino A, Stracquadanio S, Cosentino F, et al (2026)

Beyond the Usual Suspects: Emerging Pseudomonas Species in Clinical and Environmental Niches.

International journal of molecular sciences, 27(14):.

Non-aeruginosa Pseudomonas (NAP) species represent a diverse and ubiquitous group of Gram-negative bacteria inhabiting a wide range of environmental niches, from soil and water to plant rhizospheres and clinical settings. While Pseudomonas aeruginosa has historically dominated clinical and research focus, the significance of NAP species, such as Pseudomonas fluorescens, Pseudomonas putida, and Pseudomonas stutzeri, as both opportunistic human pathogens and versatile biotechnological agents is increasingly recognized. Their remarkable genomic plasticity, driven by large accessory genomes and mobile genetic elements, underpins their metabolic versatility and adaptability but also facilitates the acquisition of virulence determinants and antibiotic resistance genes, contributing to their emergence in healthcare settings, particularly among immunocompromised individuals. This review provides a comprehensive analysis of NAP species, focusing on recent advances in their taxonomy facilitated by genomic tools like Whole-Genome Sequencing (WGS) and Multilocus Sequence Typing (MLST), which reveal complex species groups and challenge traditional classifications. We delve into the genomic landscape, exploring pangenome dynamics, horizontal gene transfer (HGT), and the genomic signatures that may differentiate clinical from environmental isolates. The clinical relevance of NAPs is examined, detailing the spectrum of infections, epidemiological trends, risk factors, and insights into virulence mechanisms, including secretion systems (T3SS, T6SS) and pathogenicity islands. Addressing a critical need, this review incorporates detailed sections on the diagnostic challenges posed by NAPs, including common misidentifications and the role of modern techniques like MALDI-TOF MS and WGS, and outlines current and novel therapeutic strategies, considering the growing problem of antimicrobial resistance (AMR) within this group. Furthermore, the biotechnological applications of NAPs in bioremediation and biocatalysis are discussed alongside evolving biosafety considerations, reflecting the shift from strict containment to integrated monitoring approaches for genetically engineered strains. By synthesizing current knowledge and highlighting research gaps, this review underscores the necessity of integrated, One Health approaches to understand and manage the dual nature of non-aeruginosa Pseudomonas species as both environmental inhabitants and clinically relevant pathogens.

RevDate: 2026-07-29
CmpDate: 2026-07-29

Qiao Y, Zhang X, Jing R, et al (2026)

Molecular Epidemiology, Phenotypic and Genomic Characterization of Multidrug-Resistant Enterococcus Faecium Isolated from Bovine Mastitis in Ningxia, China (2019-2024).

Microorganisms, 14(7): pii:microorganisms14071424.

Multidrug-resistant (MDR) Enterococcus faecium is an opportunistic pathogen. Its resistance and virulence genes can spread through the food chain, posing risks to public health. This study investigated the antimicrobial resistance and genomic characteristics of MDR E. faecium isolated from milk samples from cows with mastitis in Ningxia between 2019 and 2024. From 2019 to 2024, 1341 milk samples were collected in Yinchuan, Yinnan, and Yinbei. MDR E. faecium was identified using plate screening, mass spectrometry, broth microdilution, and hemolysis detection. Whole-genome sequencing enabled SNP, MLST, pan-genome, and COG analyses, focusing on ARGs and MGEs. MRPP, AMOVA and PCoA were applied to compare gene communities and identify driver genes. Ninety-one E. faecium strains were isolated. Resistance to florfenicol, ceftiofur, and chloramphenicol exceeded 60%, while resistance to vancomycin and linezolid showed an overall increasing trend over the study period. Phylogenetic clustering revealed two subtypes, three clades, and 10 novel STs. Spearman correlation analysis revealed strong positive correlations among the resistance genes optrA, cfr(A), and vanF. Antibiotic resistance, particularly MDR, increased over time, and strains carried diverse ARGs and MGEs. Overall, strengthened surveillance of mastitis-derived E. faecium is warranted to support the control of bovine mastitis and safeguard public health.

RevDate: 2026-07-29
CmpDate: 2026-07-29

Shapovalova VV, Shaidullina ER, Ivanchik NV, et al (2026)

Genomic phylogeny, taxonomic classification, genetic landscape, and antibiotic resistance of clinical Stenotrophomonas isolates.

Frontiers in microbiology, 17:1854839.

BACKGROUND: Stenotrophomonas spp. are increasingly recognized opportunistic pathogens characterized by intrinsic multidrug resistance. Despite recent taxonomic reclassification and the delineation of new species within the S. maltophilia complex (Smc), routine diagnostics often misidentify these species, and their distinctive antibiotic resistance traits and genetic makeup remain underexplored. We aimed to characterize the phylogenomic diversity, genotypic and phenotypic antibiotic resistance traits, virulence markers and horizontally transferrable elements in a large national cohort of clinical Stenotrophomonas isolates.

METHODS: We analyzed 323 clinical isolates from the Russian sentinel surveillance program (2002-2021). Analysis included MALDI-TOF MS identification, broth microdilution susceptibility testing to six agents [aztreonam-avibactam, colistin, levofloxacin, minocycline, tigecycline, and trimethoprim-sulfamethoxazole (TMP-SMX)], and long-read Nanopore WGS with hybrid polishing. We performed average nucleotide identity (ANI)-based classification, core- and accessory-genome comparisons, and comprehensive profiling of antibiotic resistance genes (ARGs), virulence factors (VFs), plasmids, and prophages.

RESULTS: While MALDI-TOF MS identified all isolates as S. maltophilia, ANI (≥95% threshold) and core-genome phylogeny revealed 17 distinct species: 52% S. maltophilia sensu stricto, 36.8% across eight other ICNP-named Smc species, and 11.2% within eight "novel" genomic lineages identified elsewhere. Pan-genome comprised 24,815 genes (9.6% core genes), with accessory genes showing robust species-level clustering in UMAP. TMP-SMX resistance was prevalent (40.2%) among all species, though only 13.1% of resistant isolates harbored sul genes. Intrinsic β-lactamase genes (bla L1- and bla L2-like) were ubiquitous but exhibited high sequence variability (up to 20.6% and 25.1%) and evidence of inter-species exchange. Phosphoethanolamine transferase genes (mcr-5- and mcr-8-like) were identified as intrinsic chromosomal components. Acquired ARGs were primarily associated with chromosomal integrative conjugative and mobilizable elements (ICEs/IMEs). Plasmids (2.5 to >334 kb) were identified in 8% of isolates across 10 species, including a cluster of eight nearly identical replicons distributed among five species despite lacking canonical mobilization genes. Prophages were ubiquitous (median: 4 per genome), dominated by Caudoviricetes and Inoviridae. Several VFs linked to motility, capsule formation, and invasion/toxicity were universally present across the genus.

CONCLUSION: Our findings demonstrate that current routine diagnostics fail to resolve the significant taxonomic diversity within clinical Stenotrophomonas isolates, leading to widespread species misidentification. Integrating the genomic and phenotypic insights from this study into routine surveillance is essential to overcome these diagnostic limitations and to modernize the clinical management of infections caused by this formidable genus.

RevDate: 2026-07-29

Pani S, Dabbaghie F, Marschall T, et al (2026)

gaftools: a toolkit for analyzing and manipulating pangenome alignments.

Bioinformatics (Oxford, England) pii:8746777 [Epub ahead of print].

MOTIVATION: Linear reference genomes are ubiquitously used in genomics research, despite known biases associated with their use. In recent years, there has been a shift towards graph-based reference genomes to address some of these biases, which has required development of new algorithms and file formats. This has created a necessity for new tools capable of utilizing these formats and performing operations similar to those carried out by traditional methods.

RESULTS: In this paper we present "gaftools", a multi-purpose tool that introduces several utilities for processing graph alignments in GAF format. Gaftools enables users to index and sort alignments, with graph ordering serving as a necessary step for the sorting process. Additionally, it allows users to view subsets of alignments and perform realignment using the wavefront alignment algorithm, among other features. Many of these functionalities are inspired by SAMtools, which provides similar operations for linear genomes, while gaftools adapts and extends them for pangenomes.

AVAILABILITY: gaftools is available under MIT license at https://github.com/marschall-lab/gaftools.

RevDate: 2026-07-29
CmpDate: 2026-07-29

Smith OER, Clemente CM, Andreeva A, et al (2026)

Structural basis of biofilm formation mediated by the Pseudomonas aeruginosa fibrillar adhesin CdrA.

bioRxiv : the preprint server for biology pii:2026.07.13.738186.

Many bacteria, including the important human pathogen Pseudomonas aeruginosa , are naturally found in antibiotic-tolerant, multicellular biofilms. Cell-cell interactions within P. aeruginosa biofilms are mediated by a large fibrillar adhesin called CdrA in an extracellular polysaccharide-dependent manner. Here, we report an electron cryomicroscopy structure of the 60 kDa CdrA adhesive N-terminus, which combined with electron cryotomography of focused-ion beam milled specimens, allows us to derive a complete in situ model of the native adhesin. Our structure reveals a small adhesive domain (called ADEPT) at the distal tip of CdrA that is nearly perfectly conserved across the P. aeruginosa pangenome, with structural similarity to previously reported sugar-binding domains in multiple bacterial species. Inhibitory nanobodies targeting CdrA that reduce biofilm formation bind to epitopes in, or close to, the ADEPT on bacterial cells. Furthermore, structure-guided mutagenesis of residues within the ADEPT abolishes bacterial aggregation, and genomic deletion of the whole ADEPT leads to strong attenuation of biofilm formation. Our data forms a rational basis for future targeted inhibition of pathogenic P. aeruginosa biofilms and elucidates the mechanism of biofilm formation mediated by fibrillar adhesins that are widespread in bacteria.

RevDate: 2026-07-29
CmpDate: 2026-07-29

Abdurakhmonov IY (2026)

Plant RNA interference from antiviral silencing to multiplex trait engineering for climate-resilient crops.

Frontiers in plant science, 17:1871239.

RNA interference (RNAi) in plants has evolved from an unexplained antiviral and transgene interference phenomenon into a general regulatory platform for sequence-guided gene suppression, chromatin control, systemic signaling, and phenotypic plasticity. This Review synthesizes six decades of plant RNAi, tracing its progression through conceptual bottlenecks and technological solutions. Early work established that RNA-derived homology could suppress viral infection and transgene expression. Mechanistic studies then revealed a diversified plant silencing system involving Dicer-like proteins, Argonautes, RNA-dependent RNA polymerases, systemic movement, and RNA-directed DNA methylation. In parallel, RNAi moved into crop design, enabling targeted modification of yield, fiber quality, flowering, disease resistance, allergenicity, fertility, plant architecture, lignin content, nutrient composition, and pest resistance across diverse species. Importantly, RNAi is not merely a historical precursor to genome editing. It retains distinct value because it can tune gene dosage, silence multigene families, uncover compensatory network responses, and perturb upstream regulatory nodes, such as phytochrome RNAi in cotton, where partial suppression simultaneously improves several negatively correlated traits. Most recently, host-induced silencing, spray-induced dsRNA, nanocarrier delivery, and CRISPR-associated RNA tools have repositioned RNAi as a versatile breeding platform. The future lies in convergence with genome editing, using pangenome-informed, allele-aware target design and combined RNAi-editing pipelines. The lesson learned is that useful crop engineering often requires rebalancing endogenous networks rather than permanent gene knockout. In this review, the historical developmental phases are used carefully: the formal molecular term RNA interference emerged in the late 1990s, while earlier plant work on antiviral resistance, co-suppression and post-transcriptional gene silencing anticipated the same sequence-guided logic. At the same time, practical deployment remains constrained by variable knockdown, off-target risk, construct instability, environmental degradation of sprayed RNA, delivery cost, resistance evolution in target pests or pathogens, regulatory classification, and public acceptance; these constraints are discussed as platform-specific design and risk-assessment issues rather than as generic barriers.

RevDate: 2026-07-27
CmpDate: 2026-07-27

Tan H, Yang S, X Zhang (2026)

The Neurospora crassa Pangenome: A Robust Framework for Population-Scale Analysis and Structural Variant Discovery.

Journal of fungi (Basel, Switzerland), 12(7):.

Neurospora crassa is a widely distributed ascomycete with high genetic diversity, yet reliance on limited reference genomes has hindered a comprehensive understanding of its genetic landscape. To address this limitation, we integrated the functional annotation of the FGSC2225 genome with a comprehensive comparative genomic analysis of N. crassa strains. FGSC2225 gene and transposable element (TE) proportions mirrored those of FGSC2489, though TE levels were significantly higher than those in sister species Sordaria macrospora. Phylogenetic analysis resolved the N. crassa population into two primary lineages: Clade A (including FGSC2489 and FGSC2225) and Clade B (including FGSC4830), with the former exhibiting larger genome sizes. Leveraging de novo assemblies of 72 high-quality draft genomes, we constructed a comprehensive pangenome to investigate the molecular evolution of various gene families. For example, systematic phylogenetic analysis of the HET-domain-containing gene family and three stress-related families-heat shock transcription factor, basic leucine zipper, and Cytochrome P450-demonstrated varying degrees of conservation and presence/absence variation across the lineages. Addressing the limitations of current genomic resources, this work provides a pangenomic framework to detect rapid adaptive evolution in filamentous fungi. This methodology serves as a robust template for identifying transcription factors, effectors, and structural variations critical to stress response and virulence in diverse fungi.

RevDate: 2026-07-27
CmpDate: 2026-07-27

Ma R, Wang H, Wei Y, et al (2026)

Population genomics reveals gene flow and positive selection patterns in the wine-related yeast Hanseniaspora uvarum.

Stress biology, 6(1):.

Hanseniaspora uvarum is a representative non-Saccharomyces species that plays a significant role in fermentation processes such as winemaking. In recent years, this species has gained attention in food engineering and evolutionary biology. However, the population genomic signatures in this species remain poorly understood. In this study, a population genomics analysis was conducted on 151 H. uvarum strains (45 from Ningxia, China; 21 from other regions of China; 67 from Australia; and 18 from other regions or of unspecified origin), and a pangenome analysis was performed on 159 strains, incorporating eight additional genome assemblies. Phylogenetic analysis, ancestry coefficient analysis, and principal component analysis generally distinguished Chinese strains from those sampled on other continents. However, substantial post-divergence gene flow and introgression were inferred between intercontinentally paired clades. Positively selected candidate genes exhibited region-specific patterns: GO terms related to the positive regulation of filamentous growth in response to external stimuli were significantly enriched in the Ningxia strains; the stress-related GO term "cytoplasmic stress granule" was significantly enriched in both the Ningxia and Australian strains, but with distinct sets of associated genes. Although the samples were primarily isolated from anthropogenic environments, H. uvarum exhibited an open pangenome, indicating substantial adaptive potential to diverse stresses. This study advances our understanding of the evolutionary dynamics of H. uvarum and establishes a genomic foundation for future ecological and industrial research on this yeast.

RevDate: 2026-07-27

Ghareeb AM (2026)

Investigation in genomic signatures of helicobacter pylori through a microbial-genome wide association study exploring virulence variation.

Diagnostic microbiology and infectious disease, 116(3):117567 pii:S0732-8893(26)00317-2 [Epub ahead of print].

BACKGROUND: Helicobacter (H.) pylori is characterized by a high degree of genomic diversity, with regional differences in virulence determinants. This study aims to explore genomic composition, phylogeography and accessory-gene relationships of Iraqi H. pylori isolates in a global contextualized dataset.

METHODS: A total of 198 H. Pylori genomes were reviewed, including 41 isolates sequenced from gastric samples of patients undergoing diagnostic endoscopy at Al-Yarmouk Teaching Hospital, Baghdad from June 2024 to February 2025. Illumina MiSeq was used to sequence genomes, which were quality-filtered and assembled using SPAdes. Prokka was used to perform annotation and Roary to infer pan-genome structure. FastTree was used to reconstruct core-genome phylogeny. Anatomical micro-niche (corpus vs. other gastric sites) were explored with pan-genome-wide association study (pan-GWAS) with Pyseer (linear mixed model, kinship based on the core alignment).

RESULTS: The H. Pylori pan-genome showed 955 core gene families and 9,400 accessory genes. Isolates from Iraq were polyphyletic, mixing with European and Middle Eastern lineages. In the primary Pyseer linear mixed-model analysis, Benjamini-Hochberg correction across 3,259 valid lrt-pvalue tests identified 151 FDR-significant associations; after excluding rows with problematic Pyseer diagnostic notes, 70 unflagged loci remained significant. The strongest unflagged positive association was group_2036, whereas a co-occurring block including cagS, cagT, and virB4_1 was strongly depleted in corpus-derived isolates. These signals implicate accessory-genome variation in gastric micro-niche adaptation while also underscoring the need to interpret flagged Pyseer rows cautiously.

CONCLUSIONS: The genomic variation of squamous H. Pylori isolates in Iraq corresponds to the global recombination trends and to the regional admixture. The corpus sampling-related accessory-gene cluster implies possible micro-niche adaptation. These preliminary results highlight the necessity of large, stratified Middle Eastern cohorts and long-read sequencing to dispel functional genetic constructions of tissue tropism and virulence.

RevDate: 2026-07-28

Tian M, Liang Y, Lu J, et al (2026)

Global Genomic Analysis of Bovine-Associated Klebsiella pneumoniae Reveals Genetic Diversity and Resistance-Virulence Profiles.

Biology, 15(14): pii:biology15141215.

Bovine-associated Klebsiella pneumoniae is an important bacterial species linking animal health, microbial ecology, and One Health-oriented antimicrobial resistance research. In this study, we performed a global genomic analysis of 1291 publicly available bovine-associated K. pneumoniae genomes collected from 18 countries between 2005 and 2024 using data retrieved from NCBI. MLST, core-genome phylogenetic analysis, pangenome analysis, CARD, VFDB, and PlasmidFinder were used to characterize sequence types, genomic diversity, antimicrobial resistance-associated genes, virulence-associated genes, and plasmid replicons. A total of 256 sequence types were identified, among which ST107 was the most common. Core-genome phylogenetic analysis revealed multiple genomic lineages, while pangenome analysis identified 46,325 gene clusters, including 1967 core genes and 40,595 cloud genes, indicating an open pangenome structure and substantial accessory gene diversity. Virulence-associated genes were unevenly distributed, with yagZ/ecpA being the most frequently detected determinant. In total, 138 antimicrobial resistance-associated genes or potential resistance determinants were detected across 16 antimicrobial categories, including clinically important β-lactamase- and carbapenemase-associated genes. IncF-family plasmid replicons, particularly IncFIB(K)_1_Kpn3, were frequently detected, suggesting widespread plasmid replicon-associated genomic backgrounds; however, physical co-localization between resistance genes and specific plasmid backbones could not be confirmed. Overall, this study reveals the genetic diversity, resistance-associated gene reservoir potential, heterogeneity of virulence-associated genes, and plasmid replicon backgrounds of bovine-associated K. pneumoniae. Importantly, the genome-predicted AMR potential identified in this study should not be interpreted as confirmed phenotypic resistance without further experimental validation. These findings provide genomic insights for risk surveillance, candidate control-target screening, and microbiota-oriented intervention research.

RevDate: 2026-07-28

Meroni G, Sora VM, Soggiu A, et al (2026)

Comparative Genomic Analysis of Bovine and Publicly Available Human Streptococcus agalactiae Genomes.

Animals : an open access journal from MDPI, 16(14): pii:ani16142257.

BACKGROUND/OBJECTIVES: Streptococcus agalactiae is one of the most significant pathogens causing infections in humans and mastitis in dairy cattle. This work focused on a comprehensive comparative pan-genomic analysis of bovine and human Group B Streptococcus to elucidate the genetic mechanisms underlying host adaptation and dissemination.

METHODS: Isolates of S. agalactiae from quarter milk samples from dairy herds in Lombardy (Italy), along with human strains, were considered. Whole genome sequencing was used to compare core and accessory genomes, assign sequence types, and find virulence and resistance factors.

RESULTS: 30 sequence types were detected, of which two (ST12 and ST23) in common between bovine and human. The allele frequencies for resistance determinants revealed elevated rates for tetM (59.8% overall, 66.5% in human clinical isolates), ermB (17.3% overall, 20.4% in human clinical isolates), and ant(6)la (8.9% overall, 12.8% in human clinical isolates). Bovine strains had accessory gene clusters linked to lactose metabolism and immunological evasion, whereas human isolates were concentrated in regions related to adhesion and antibiotic resistance.

CONCLUSIONS: Comparative pan-genomics show that there is a small genetic overlap between bovine and human Group B Streptococcus populations.

RevDate: 2026-07-26
CmpDate: 2026-07-24

Zhang C, Liu C, Hua J, et al (2026)

The Pan-Genome of 1107 Spotted Sea Bass Accessions Reveals Gene Evolution Patterns During Domestication Selection.

Molecular ecology resources, 26(5):e70180.

Spotted sea bass (Lateolabrax maculatus) is an ecologically and economically important species that supports a large-scale aquaculture industry in China. However, most farmed populations have not undergone systematic breeding programs and are characterised by undocumented empirical selection. Understanding genetic diversity dynamics and gene evolution patterns for farmed populations is critical for designing appropriate breeding strategies. Here, we assembled a chromosome-level reference genome using HiFi and Hi-C sequencing technologies, resulting in a highly contiguous assembly with a total size of 636.65 Mb and a scaffold N50 of 27.27 Mb. Furthermore, we performed population genomic analyses using 1107 representative wild and farmed individuals. Significant genetic differentiation and reduced genetic diversity were observed in farmed populations relative to wild populations. Pan-genome of 1107 accessions further captured 86.48 Mb of non-reference novel sequences (NRNS) and 206 novel genes absent from the reference genome. Gene presence/absence variation (PAV) analysis revealed that farmed populations harbour fewer genes than wild counterparts, implying gene loss and negative selection during the domestication process. Lost genes were enriched in functions related to immune response and nervous system development, highlighting wild populations as important genetic resources for future breeding programs. Notably, selection signatures and comparative PAV analyses across independent farmed populations revealed parallel domestication effects at both SNP and gene PAV levels, suggesting similar selective pressures despite varied aquaculture practices and domestication histories. Our study provides the first pan-genome of spotted sea bass that deciphers genetic diversity dynamics and gene evolution patterns during domestication selection and establishes valuable genomic resources for sustainable genetic improvement of this species.

RevDate: 2026-07-27
CmpDate: 2026-07-24

de Magalhães PC, Felice AG, S de Castro Soares (2026)

Pangenomic and genomic plasticity analyses of the genus Rickettsia.

Brazilian journal of microbiology : [publication of the Brazilian Society for Microbiology], 57(1):.

The Rickettsia genus comprises obligate intracellular bacteria transmitted by arthropods and responsible for clinically relevant zoonoses, rickettsioses, such as spotted fever and typhus. The difficulty of cultivating these bacteria in vitro reinforces the importance of in silico approaches, such as pangenomic and genomic plasticity analyses. This study analyzed 165 genomes from 31 Rickettsia species available in the REFSEQ (NCBI) database. Tools such as Orthofinder, ANIclustermap, Gegenees, and Mauve were used to classify genes into core, shared, and singletons, assess genomic similarity, and identify structural rearrangements. The results indicate that the genus has an open pangenome (α = 0,842), suggesting high genetic variability and adaptive and expansion potential. Species such as R. typhi exhibited a nearly closed pangenome (α = 0,999), with high genomic conservation, whereas R. rhipicephali showed an open pangenome (α = 0,876), reflecting greater plasticity and intraspecies diversity. Functional categorization of genes revealed that the core genome is associated with vital functions, while singletons include genes related to genetic mobility, indicating possible acquisition through horizontal transfer. Synteny analysis demonstrated high gene conservation in R. typhi and extensive structural reorganization in R. rhipicephali. Statistical correlation reinforced the stability of the core genome regardless of pangenome expansion and revealed an inverse relationship between the number of singletons and the value of α. The findings demonstrate the existence of contrasting evolutionary trajectories within the Rickettsia genus, with conserved, specialized species coexisting alongside genetically dynamic species that are more adaptable to different niches. Thus, this study expands the understanding of clonality, genomic plasticity, and functional diversity within the genus, providing support for future investigations into virulence factors, vaccine targets, and bacterial evolution.

RevDate: 2026-07-25

Nagy GR, Munkácsy G, Murmu A, et al (2026)

Beyond the linear genome: how reference bias threatens preventive medicine and geroscience.

GeroScience [Epub ahead of print].

Genotype-first screening compares patient genomes to a standard reference. The GRCh38 linear assembly contains reference minor alleles (RMAs), posing a critical structural vulnerability for automated clinical variant classification. To determine whether RMAs systematically generate false-positive annotations, we conducted an observational case series of 20 healthy adults undergoing preventive genome sequencing. Automated bioinformatic processing was performed using the GRCh38 linear reference to identify "high-impact" variant annotations generated at loci where GRCh38 differs from the population consensus major allele. Among all 20 participants (100%), linear alignment to GRCh38 systematically misclassified functional major alleles as false-positive "high-impact" annotations at three distinct loci (SLC37A4, CIMIP2A, and GPR33). These artifacts occurred solely because analysis software mathematically defined the healthy wild-type state as a deviation from the rare RMA. Cross-referencing with gnomAD confirmed these as population-dominant benign alleles. Furthermore, this inherent reference bias introduces a theoretical risk of false-negative classifications if RMAs mathematically mask true pathogenic variants. Automated pipelines using linear references systematically misclassify healthy alleles as high-impact functional annotations. While downstream population-frequency filters manage false-positive artifacts, this retrospective patching creates an unsustainable bottleneck for population-scale screening. Transitioning to graph-based pangenome references represents a highly promising approach to directly resolve these diagnostic vulnerabilities at the alignment level, though computational and standardization challenges must first be addressed to ensure the accuracy of large-scale aging research.

RevDate: 2026-07-27

Jiménez-Blanco A, López-Villellas L, Moure JC, et al (2026)

Theseus: Fast and Optimal Affine-Gap Sequence-to-Graph Alignment.

Bioinformatics (Oxford, England) pii:8742314 [Epub ahead of print].

MOTIVATION: Sequence-to-graph alignment is a central problem in bioinformatics, with applications in multiple sequence alignment (MSA) and pangenome analysis, among others. However, current algorithms for optimal affine-gap alignment impose high memory and computational requirements, limiting their scalability to aligning long sequences to complex graphs. Practical solutions partially address this problem using heuristic strategies that ultimately trade off optimality for speed.

RESULTS: This work presents Theseus, a novel, fast, and optimal affine-gap sequence-to-graph alignment algorithm. Theseus leverages similarities between genomic sequences to accelerate the alignment computation and reduces the overall memory requirements without compromising optimality. To that end, Theseus processes only a subset of the dynamic programming cells, using a sparse-data strategy that enables efficient sequence-to-graph alignment. Moreover, our algorithm supports optimal affine-gap alignment on arbitrary directed graphs, including those with cycles. We evaluate Theseus on two key problems: multiple sequence alignment (MSA) and pangenome read mapping. For MSA, we compare it against SPOA, abPOA, and POASTA. Theseus is 1.6× to 17.6× faster than POASTA, and 7.3× faster, on average, than SPOA, both optimal aligners. Compared with abPOA, Theseus ensures optimality and scales to the largest problems. For pangenome read mapping, we benchmark Theseus against the alignment stage of the mapping tool vg map, along with the alignment kernels of SPOA, abPOA, and POASTA. Theseus outperforms the other methods, showing a 1.9× to 16.9× speedup on short reads. Moreover, Theseus is 1.5× to 36.3× faster than vg when aligning against synthetic cyclic graphs.

AVAILABILITY: Theseus code and documentation are publicly available at https://github.com/albertjimenezbl/theseus-lib.

RevDate: 2026-07-24
CmpDate: 2026-07-24

Eripogu KK, Maharathi P, WH Li (2026)

Comparative genomics of Nocardia seriolae reveals a conserved metabolic core and extensive accessory genome plasticity.

Frontiers in microbiology, 17:1881314.

INTRODUCTION: Fish nocardiosis is a chronic and economically significant bacterial disease in aquaculture, yet its genomic basis remains poorly resolved beyond single-species studies. It remains unclear whether fish-associated Nocardia share conserved persistence-associated features or exhibit lineage-specific genomic diversification.

MATERIALS AND METHODS: We conducted a comparative genomic analysis of 22 Nocardia genomes, including 20 N. seriolae isolates and single representatives of N. salmonicida and N. crassostreae. Genome-wide analyses included phylogenomics, gene-content comparison, pangenome analysis, functional annotation, virulence-associated homolog screening, genomic island detection, and secondary biosynthetic gene cluster prediction.

RESULTS: The conserved genome core was enriched in central metabolism, lipid-associated cell envelope biogenesis, iron acquisition, and stress-response pathways. Virulence-associated homologs were dominated by persistence-associated and metabolic functions, whereas classical toxin systems were limited, although several transport- and secretion-associated homologs were detected, consistent with their potential contribution to host interaction and intracellular persistence. Phylogenomic and gene-content analyses revealed clear species-level divergence but limited host-associated structuring within N. seriolae. Pangenome analysis supported a robust open pangenome structure (γ = 0.386), with extensive accessory gene diversity enriched in regulatory functions, mobile genetic elements, and secondary metabolic pathways. Genomic islands were dominated by insertion-sequence-associated genes, recombinases, regulators, and hypothetical proteins, whereas prophage- and toxin-related signatures were rare. Secondary metabolite analysis revealed extensive biosynthetic diversity, with most biosynthetic gene clusters showing low similarity to characterized reference pathways. However, ectoine- and nocobactin-associated pathways were broadly conserved.

CONCLUSION: These genome findings are consistent with a persistence-associated pathogenicity model in which fish-associated Nocardia, particularly N. seriolae, may depend more on metabolic resilience, stress adaptation, iron acquisition, and accessory genome plasticity than on classical toxin-mediated virulence. Collectively, the results highlight the importance of accessory genome diversification, iron acquisition, and stress adaptation in shaping host-associated lifestyles and provide a comparative genomic foundation for future functional investigations and aquaculture disease-management strategies.

RevDate: 2026-07-23

Whelan FJ (2026)

How the social lives of bacteria affect their pangenome.

Essays in biochemistry pii:237850 [Epub ahead of print].

Although the study of microbes started with type strains and reference genomes, advances in sequencing technology and new interest in mixed microbial communities have made us aware that a single genome cannot and does not reflect the diversity of a given bacterial species. Bacteria rarely occupy an environmental or host niche alone and quickly diversify into strains upon colonization of a new niche. The genetic diversity present within a phylogenetically related set of bacterial strains (the 'pangenome') is influenced by the niche that they occupy and how they interact with the other microorganisms that they share that niche with. In this review, I examine how the social lives of bacteria can affect their genetic diversity and the bioinformatic techniques that we use to detect that diversity.

RevDate: 2026-07-23

Lao J, Zhang L, Huang X, et al (2026)

Longitudinal surveillance of antibiotic resistance and virulence evolution in Clostridioides difficile: a 4-year retrospective study of hospitalized patients in a tertiary hospital in China.

Microbiology spectrum [Epub ahead of print].

UNLABELLED: Clostridioides difficile (C. difficile) is the primary pathogen responsible for nosocomial infectious diarrhea and pseudomembranous colitis. In China, metronidazole and vancomycin are the preferred treatments for C. difficile infection (CDI). This study aimed to investigate the evolution of vancomycin (VA) and metronidazole (MTZ) resistance, as well as the longitudinal changes in virulence over time, using next-generation sequencing, drug susceptibility tests, and analysis of resistance and virulence genes. Additionally, we monitored the emergence of the highly virulent C. difficile strain RT027 and the spread and potential outbreak of C. difficile in the hospital setting. A random stratified sampling method was used to select 114 fecal samples from inpatients at Affiliated Hangzhou First People's Hospital, School of Medicine, Westlake University, between 2021 and 2024. Clinical data from the enrolled patients were also collected. We conducted antigen and toxin protein detection for C. difficile, strain isolation and identification, drug sensitivity tests, whole genome sequencing, and bioinformatics analysis. This included comparisons of drug resistance genes, detection of toxin genes, and the construction of phylogenetic trees based on pan-genome analysis to investigate the resistance and toxin gene variations in C. difficile. Among the 114 samples collected from Affiliated Hangzhou First People's Hospital, School of Medicine, Westlake University, no vancomycin- or metronidazole-resistant strains were identified. However, the average minimum inhibitory concentration (MIC) of C. difficile to vancomycin increased annually (H = 33.208, P < 0.05). The average MIC of C. difficile to metronidazole was highest in 2022 but decreased in 2023 and 2024 (H = 41.990, P < 0.05). Notably, in 2024, one C. difficile strain exhibited an MIC for metronidazole at the resistance threshold (2.00 μg/mL). Further Spearman correlation analysis of the strain years with drug sensitivity results revealed a positive correlation between strain years and the MIC levels of vancomycin and metronidazole (r = 0.528, P < 0.05; r = 0.377, P < 0.05). The proportion of toxin-producing strains increased annually, with 100% of strains in 2024 producing toxins, representing the highest proportion compared to the previous three years (X[2] =11.75, P < 0.05). Both vancomycin and metronidazole remain effective for the treatment of CDI in clinical practice. However, the sensitivity of C. difficile to these two drugs is gradually decreasing, and the rate of toxin gene carriage is also rising in clinical cases. No hospital outbreaks of C. difficile infections were identified in this study.

IMPORTANCE: Clostridioides difficile has developed resistance to multiple antibiotics, including cephalosporins, clindamycin, and fluoroquinolones. This has exacerbated the global antibiotic resistance crisis. In China, according to current treatment guidelines, vancomycin and metronidazole are the preferred first-line drugs for treating C. difficile infections. However, there are reports indicating the emergence of new resistance to both vancomycin and metronidazole. Although there is extensive research on the long-term antibiotic resistance of C. difficile abroad, research on the continuous monitoring of antibiotic resistance and potential outbreaks of C. difficile in China is relatively limited. To fill this gap, we studied positive C. difficile strains from a tertiary general hospital in China. Through Next-Generation Sequencing (NGS), drug sensitivity testing, and analysis of drug resistance and virulence genes, we revealed the evolution of C. difficile's resistance to vancomycin and metronidazole, as well as changes in virulence, and monitored the spread within the hospital and potential outbreaks of C. difficile.

RevDate: 2026-07-23
CmpDate: 2026-07-23

Liu Q, Lian J, H Tang (2026)

Evolutionary and pan-genomic analysis of the bZIP gene family in 21 Camellia sinensis.

Frontiers in plant science, 17:1885680.

Basic leucine zipper (bZIP) transcription factors are important regulators of plant development and stress responses, yet their evolutionary dynamics in tea plant have largely been inferred from a single reference genome. Here, we performed a broad evolutionary and pan-genomic analysis of the bZIP family using 1,015 plant genomes and 21 Camellia sinensis genomes. Across plants, 81,340 bZIP genes were identified, revealing broad conservation of this family across major lineages and a significant copy-number expansion in angiosperms. In tea, 1,635 non-redundant bZIP genes were identified, with 73-88 members per genome, indicating an overall conserved family size among tea germplasms. Phylogenetic analysis classified these genes into 13 subfamilies, among which S, A, D, I and G represented the major expanded groups. Orthogroup analysis resolved 77 bZIP orthogroups, including 22 core orthogroups and 55 dispensable orthogroups, suggesting substantial hidden variation despite stable total gene numbers. WGD/segmental duplication was the dominant expansion mechanism, accounting for 59.20% of tea bZIP genes, followed by dispersed duplication. Copy-number variation was widespread, with 72 of 77 orthogroups showing CNV across genomes. Most homologous gene pairs evolved under purifying selection, whereas dispensable genes exhibited relatively relaxed constraints compared with core genes. Transcriptome analysis in 'Shuchazao' revealed tissue-biased expression and divergent drought responses, with A, S, I and M subfamilies showing stronger PEG-induced responsiveness. Together, these results establish a pan-genome-informed bZIP resource and highlight CNV, duplication mode and dispensable gene variation as potential drivers of tea bZIP diversification.

RevDate: 2026-07-24

McKindles K, Seto K, Ahrendt S, et al (2026)

Single-cell genomics, metagenomics, and transcriptomics of Rhizophydium megarrhizum, an obligate fungal parasite of Planktothrix agardhii.

Aquatic ecology, 60(3):92.

UNLABELLED: Chytrids (phylum Chytridiomycota) are zoosporic fungi that play key roles as parasites of aquatic microorganisms, yet they are understudied and genomic resources for algal-infecting chytrids remain scarce. Here, we present the first comparative genomic analysis of multiple isolates of a single chytrid species (order Rhizophydiales) infecting the cyanobacterium Planktothrix agardhii. Isolates were collected from Sandusky Bay, Lake Erie, across two bloom years (2018 and 2019). Using single cell sequencing and metagenomic assembly, we generated individual genomes averaging 15.36 ± 0.12 Mbp in size with ~ 75% completeness, and a pangenome. Gene ontology analyses highlighted the presence of categories related to cellular structure, biosynthetic regulation, and interspecies interactions. As a preliminary exploration of gene expression during infection, we also performed RNA sequencing on a subset of size-sorted samples. These data suggest that chytrids consistently express high levels of cytoskeletal genes, alongside numerous hypothetical proteins, and that zoospores may upregulate carbohydrate-binding proteins implicated in host recognition. On the host side, P. agardhii showed transcriptional shifts in pathways associated with buoyancy and nutrient acquisition, patterns that could represent defensive adjustments or parasite-driven manipulation. Together, this study generates reference genomes for Planktothrix-infective chytrids, identifies conserved gene content across isolates from different bloom years, and provides preliminary transcriptomic insights into parasite and host responses. These resources lay the foundation for deeper investigations into chytrid genome evolution, infection biology, and their ecological roles in shaping cyanobacterial bloom dynamics.

SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at 10.1007/s10452-026-10329-8.

RevDate: 2026-07-24
CmpDate: 2026-07-24

Sebastian PJ, Schlesener C, Byrne BA, et al (2026)

Beyond AMR and virulence databases: genome wide associations of closely related Vibrio alginolyticus and emerging Vibrio diabolicus provide framework for identifying novel genetic markers.

Frontiers in microbiology, 17:1796882.

Vibrio alginolyticus is a frequently implicated species for vibriosis in humans and diverse wildlife, but it has previously been difficult to identify from the closely related and emerging Vibrio diabolicus. Comparisons of both species, including antimicrobial resistance (AMR) and virulence characterizations, are scarce and impeded by intraspecies diversity, minimal genomes, discordant classification methods, and gene databases with limited utility to understudied species. The species identities of 3,442 public domain genomes (SRA files) within the Harveyi clade were re-evaluated using genomic methods. Public genomes identified as V. diabolicus and V. alginolyticus were combined with previously published genomes isolated from humans, sea otters (Enydra lutris), or coastal environments (V. diabolicus n = 88, V. alginolyticus n = 163, Vibrio parahaemolyticus n = 287) for pangenome-wide association studies to identify species-specific gene clusters (95% identification threshold). Additional genome wide associations with isolation source (humans versus sea otters) were investigated, including AMR and virulence related gene clusters. Genomic reclassification identified 29 of 150 misclassified public domain V. alginolyticus genomes, including 26 reclassified as V. diabolicus. In total, 28 previously misclassified V. diabolicus genomes (n = 37 total) were identified, including 10 human-derived strains. GWAS identified 643 and 477 gene clusters specific to V. alginolyticus and V. diabolicus, respectively, while some multilocus sequencing analysis (MLSA) gene clusters were non-specific. Gene clusters (n = 109) associated with either V. alginolyticus isolated from humans or sea otters were identified including one annotated to a multidrug resistance gene (mdtk_1). No V. diabolicus gene clusters were associated with host species after multiple comparison correction, although pre-correction associations related to antimicrobial resistance were detected (cat_1, ampC). The genomic methods of classification presented provide accurate species identification for V. diabolicus and V. alginolyticus beyond current MLSA/MLST schemes, although target species-specific genes were identified that may be useful for improved future schemes. While limited sample size of V. diabolicus hampered the ability to detect host associated markers, the GWAS approach employed provide a reusable framework for discovering insights into host adaptation and prioritizing target genes for future functional AMR and virulence validation experiments in both species.

RevDate: 2026-07-21
CmpDate: 2026-07-21

Awuah D, Hounkpe A, Anane-Asamoah J, et al (2026)

The use of artificial intelligence in advancing molecular biology in Africa: a narrative review.

Molecular genetics and genomics : MGG, 301(1):.

Artificial intelligence (AI) is rapidly becoming a core methodological pillar of molecular biology and precision medicine, and Africa is a uniquely consequential setting for this transition because the continent combines the world's greatest human genomic diversity with the most severe underrepresentation of that diversity in the datasets and reference resources on which AI models are built and benchmarked. This narrative review examines, for a genetics and genomics readership, where AI-driven methods are already strengthening African molecular biology, where the supporting evidence remains preliminary, and what is required to translate technical capability into scientifically robust and equitable benefit. The central argument is that AI is especially consequential in African molecular biology, not simply because it automates analysis, but because it can help unlock insight from African genomic diversity, pathogen biology, and clinically relevant multi-omics data that remain underrepresented in global models. Across core molecular domains, AI is accelerating protein structure prediction, high-throughput variant calling and pan-genomic reference construction, genome-wide association analysis, transcriptomic interpretation, drug discovery, and CRISPR guide design. African initiatives such as H3Africa, the African Genome Variation Project, H3ABioNet, and the H3D Centre show that locally generated datasets and African-led computational pipelines can already support meaningful discovery, from improved variant interpretation to structure-guided therapeutic prioritization. At the same time, persistent barriers remain, including underrepresentation of African genomes in training data and reference genomes, uneven computational infrastructure, limited interdisciplinary training, fragmented governance, and the risk that AI-derived benefits will remain inaccessible to the populations whose data enable them. We conclude that the future impact of AI in African molecular biology will depend less on adopting global tools in the abstract and more on building African-led datasets, validation pipelines, governance frameworks, and translational pathways that make molecular discovery both scientifically robust and equitably useful. Looking ahead, the central perspective offered by this review is that Africa's exceptional genomic diversity should be treated as a scientific asset rather than an analytical liability: realising this will require population-representative pan-genome references, sustained computational capacity, and governance structures that ensure African populations are not only the source of the underlying data but also the principal beneficiaries of the discoveries it enables.

RevDate: 2026-07-22
CmpDate: 2026-07-22

Biswas R, Sinha SS, Roy A, et al (2026)

An integrated subtractive genomics and immunoinformatics approach for designing a universal multi-epitope vaccine against Brucella spp.

Frontiers in bioinformatics, 6:1818265.

INTRODUCTION: Brucella spp. are Gram-negative bacteria accountable for brucellosis in immunocompromised individuals and livestock. Due to the slow-growing latent phenotype, current antibiotics are insufficient to treat the infection. The lack of an approved vaccine for human use against this pathogen represents a significant public health concern and indicates the urgent need for novel prophylactic interventions.

METHODOLOGY: In this study, the reverse vaccinology method was combined with pan-genome analysis to identify potential vaccine targets. Proteins have been screened for antigenicity, solubility, immunogenicity, and subcellular localization. B cell and T cell epitopes exhibiting high immunogenicity and solubility have been identified. Multi-epitope vaccine constructs have been evaluated and further analyzed depending on their physicochemical properties. Molecular docking, conformational dynamics, in silico cloning, and immune simulations were conducted to identify the optimal vaccine candidate.

RESULTS: Four proteins, trigger factor, outer membrane protein assembly factor BamA, urease subunit beta (UreB), and urease subunit alpha (UreC1) were considered for potential vaccine targets. A total of 26 B cell and 97 T cell epitopes with notable immunogenicity and solubility have been shortlisted. Twelve multi-epitope vaccine constructs were generated, among which Vc7 has been chosen based on structural and physicochemical properties. Molecular docking analysis revealed a good correlation with 2FSE and 2Z65, which were further analyzed to reveal that Vc7 exhibited stronger binding affinity (-135.24 kcal/mol) towards 2FSE, mediated by hydrophobic contacts, salt bridges, and intermolecular hydrogen bonds, making it the ideal vaccine complex and validated through a 150 ns molecular dynamics simulation. In silico cloning established construct compatibility, and immune simulation confirmed Vc7's potential to elicit T cell, B cell, antibody, and cytokine-mediated responses.

CONCLUSION: Vc7 has been identified as a structurally stable and highly immunogenic construct, suggesting its potential as a universal multi-epitope vaccine candidate for the prevention of brucellosis.

RevDate: 2026-07-22

Yu M, Wang Y, Jiang L, et al (2026)

Genomic identification, pangenome analysis, and antimicrobial resistance of Aeromonas spp. isolated from food and foodborne outbreaks.

International journal of food microbiology, 460:111979 pii:S0168-1605(26)00360-0 [Epub ahead of print].

Aeromonas spp. are extensively distributed across diverse aquatic environments and recognized as pathogens capable of causing diseases in aquatic animals. Pathogenic Aeromonas causes foodborne gastroenteritis in humans and can also lead to extra-intestinal infections. However, accurate identification of Aeromonas species remains challenging. This study aimed to accurately identify Aeromonas spp. and compare their virulence gene profiles, antimicrobial resistance patterns, and molecular evolutionary relationships. A total of 42 Aeromonas isolates were obtained from retail food and foodborne disease outbreaks. They were initially identified using matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) and further confirmed by genomic methods. Average nucleotide identity (ANI) can accurately identify Aeromonas species. However, a higher ANI threshold is required to distinguish closely related species. The genus Aeromonas was found to possess an open pan-genome, enabling the acquisition of new genetic elements and enhancing environmental adaptability. All isolates encoded β-lactamase resistance genes, and 90.5% (38/42) of these were conferred resistance to ampicillin and amoxicillin-sulbactam, with the 95% confidence interval (CI) of 77.9%-96.2%. Some strains harbored antimicrobial resistance genes, such as mcr, tetE, sul, qnr, and so forth, and conferred resistance to the corresponding antibiotics. Some strains contained mobile elements carrying antimicrobial resistance gene clusters, such as transposon Tn5393 and antibiotic-resistant plasmids, providing mechanistic insights into their potential for horizontal antimicrobial gene transfer and adaptive evolution. Certain Aeromonas species possessed numerous virulence genes, including ast, hlyA, rtx, aerA, and hutX, and genes encoding flagellar, pili, and secretion systems. A. dhakensis, A. salmonicida, A. hydrophila, A. veronii, and A. enteropelogenes were predicted to have higher virulence potential. In contrast, A. caviae, the main Aeromonas species associated with foodborne disease outbreaks, exhibited relatively fewer virulence genes. This study emphasized the pathogenic potential and antimicrobial resistance profiles of Aeromonas species. Continuous monitoring of resistance patterns and contamination levels in food products is crucial for minimizing infection risks and preventing disease outbreaks caused by Aeromonas spp.

RevDate: 2026-07-22
CmpDate: 2026-07-22

Tran TTH, Hoang TH, Tran MH, et al (2026)

VN1K is a pangenome-informed multi-omics and phenomics resource for the Vietnamese population.

Nature communications, 17(1):.

The population of Vietnam remains underrepresented in global genomic databases. Here, we present VN1K, a resource of multi-omics and phenotypic information for 1011 unrelated Vietnamese individuals. We present high-depth short-read whole-genome sequencing data for all samples along with various -omics datasets. Using a high-sensitivity variant detection pipeline, which includes a pangenome graph reference and a deep-learning framework, we identify approximately 42 million variants with 7 million short insertions/deletions and 90 thousand structural variants. VN1K also features a whole-genome methylation profile based on long read sequencing. We create a genotype imputation panel with high accuracy on the Vietnamese population, allowing us to identify variants with significantly different allele frequencies in the Vietnamese population compared to other populations. We establish the functional relevance of some of these variants, particularly those in genes associated with genetic disorders, immune diseases, and drug responses, by integrating the allele frequency differences with known genotype-phenotype associations and clinical annotations. Further, we map various loci related to hepatitis B virus infection, triglyceride levels, LDL-C levels, serum glucose levels, HbA1c levels, and levels of two liver enzymes (ALT and AST). The VN1K dataset is accessible via genome.vinbigdata.org, an integrated platform with both linear and graph-based genome browsers.

RevDate: 2026-07-22

Bian J, Yang G, Xu D, et al (2026)

A pangenome of tetraploid wheat reveals the genetic architecture underlying domestication and genomic diversity for breeding.

Nature genetics [Epub ahead of print].

Tetraploid wheat (Triticum turgidum L., BBAA), a key pasta crop, serves as an untapped genetic resource with rich genomic diversity for hexaploid bread wheat improvement. Here we de novo assembled 12 genomes spanning all 10 recognized tetraploid wheat (genome BBAA) subspecies, and a graph-based pangenome was constructed. Chromosome rearrangements drove subgenome asymmetry and shaped genomic divergence, with an average of 0.25 million structural variations per accession, predominantly attributable to transposon activity. Using 736 globally distributed tetraploid wheat accessions, we identified locally adapted subgroups with untapped breeding potential and discovered a novel retrotransposon‑induced loss‑of‑function Btr1-A allele responsible for convergent adaptation of non-brittle rachis. Genome-wide association studies identified 287 loci associated with 32 traits. A homeodomain-leucine zipper transcription factor HAT14-B that enhances both spikelet number and grain size was identified. This subspecies-wide pangenome enriches Triticeae AB subgenome resources and facilitates the discovery and application of agronomically important genetic variations.

RevDate: 2026-07-22

Anonymous (2026)

A tetraploid wheat pangenome reveals genomic variation and domestication footprints.

Nature genetics [Epub ahead of print].

RevDate: 2026-07-22
CmpDate: 2026-07-22

Alfiky A, de la Rosa JMO, M Sadek (2026)

A dual-tier plasmid network model underpins the evolutionary success of pandemic Klebsiella pneumoniae ST11.

Scientific reports, 16(1):.

The convergence of antimicrobial resistance and hypervirulence in high-risk Klebsiella pneumoniae clones represents a major public health threat. However, evolutionary mechanisms enabling specific lineages to achieve pandemic dominance remain unclear. In this study, we integrated pangenomics and network analysis across 1,010 complete genomes from 38 countries. Species-wide dynamics revealed an extremely open pangenome (α = 0.59). In contrast, the dominant ST11 lineage, representing 30% of isolates, exhibited extremely low within-lineage phylogenetic diversity, consistent with a recent clonal expansion concentrated in East Asia. The East Asian ST11 lineage exhibited the lowest pangenome diversity (α = 0.86) associated with fixation of persistence and plasmid-stabilization systems and purging of redundant defense mechanisms. This configuration sustains a dual-tier plasmid network comprising a lineage-anchored IncFII(pHN7A8) replicon for vertical stability alongside high-connectivity hubs such as IncFIB(K) facilitating horizontal gene transfer. Chromosomal integration and tandem amplification of key resistance determinants (blaKPC-2, blaCTX-M-15) further reinforced this architecture. Consequently, 34.2% of isolates exhibited convergence of carbapenem resistance and hypervirulence. Within the East Asian ST11 clade, two dominant sub-lineages emerged: KL47:O13 (25.5%) and KL64:O2α (72%). Despite lower IncFII(pHN7A8) penetrance, KL64 became the dominant sub-lineage, indicating that factors beyond plasmid carriage, possibly including surface antigen properties, contribute to its epidemiological success. These findings indicate that ST11 success arises from synergy between species-wide pangenome openness and lineage-specific genomic optimization, and highlight plasmid network topology as a complementary framework for genomic surveillance of adaptive clonal expansion.

RevDate: 2026-07-23
CmpDate: 2026-07-23

Elsakhawy OK, Abouelkhair MA, SA Kania (2026)

Whole genome sequencing and molecular characterization of two Bacillus licheniformis strains isolated from hot springs in yellowstone ational park.

Frontiers in bioinformatics, 6:1867986.

Bacillus licheniformis is a Gram-positive, endospore-forming bacterium with broad biotechnological applications. Thermophilic environments such as hot springs may harbor strains with unique biosynthetic capabilities relevant to drug discovery. In this study, we isolated two B. licheniformis strains (S3 and S4) from the Five Sisters hot spring in Yellowstone National Park (68 °C and 65 °C, pH 8) and performed whole-genome sequencing using both the Oxford Nanopore long read and Illumina platforms. Hybrid de novo assembly using Unicycler yielded genome sizes of 4.80 Mbp (S3, 14 contigs) and 4.79 Mbp (S4, 22 contigs); GC contents were 45.12% and 45.10%, and N50 values were 4,546,802 bp and 2,415,736 bp, for S3 and S4, respectively. Both strains were assigned to Multi-Locus Sequence Typing sequence type ST-42. Pangenome comparison with 61 complete B. licheniformis genomes revealed an open pangenome of 10,374 genes, with 3,272 core genes, 430 soft core, 1,250 shell, and 5,422 cloud genes. AMRFinderPlus identified the blaP, encoding a class A beta-lactamase and its regulatory elements (blaI and blaR1); erm(D), encoding a 23S rRNA methyltransferase conferring macrolide-lincosamide-streptogramin B resistance; and catA, encoding a chloramphenicol O-acetyltransferase that inactivates chloramphenicol through acetylation in both strains. A chromosomal arsBC locus was identified in both B. licheniformis S3 and S4, consistent with the arsenic-rich geothermal environment of Five Sisters hot spring. These findings highlight the biosynthetic potential of B. licheniformis strains isolated from extreme environments and provide a genomic foundation for future exploration of novel bioactive compounds with potential applications in drug discovery, agriculture, and biotechnology.

RevDate: 2026-07-23
CmpDate: 2026-07-23

Lagad RR, Rafi S, A Goswami (2026)

Genomic-island cassette architecture provides interpretable signal for exploratory classification of poultry-associated Enterococcus cecorum lineages.

Frontiers in microbiology, 17:1882753.

BACKGROUND: Enterococcus cecorum is an emerging poultry pathogen whose antimicrobial resistance and host-associated traits are often carried on genomic islands. Standard comparative genomics workflows usually reduce genomes to unordered gene inventories and may miss informative neighborhood structure within island-associated modules.

METHODS: We tested whether GI (genomic island)-anchored cassette organization provides signal for distinguishing pathogenic from commensal poultry-associated E. cecorum lineages. We encoded genomic-island-anchored cassette organization as 84 genome-level summary features and evaluated this representation in 145 genomes (95 commensal, 50 pathogenic) using locked 5-fold genome-grouped cross-validation.

RESULTS: The cassette-summary Random Forest model achieved an area under the receiver operating characteristic curve (AUROC) of 0.918 ± 0.067, outperforming GI burden (AUROC 0.791 ± 0.050) and assembly-quality (AUROC 0.743 ± 0.015) baselines and performing similarly to a corrected AMR gene-content baseline (AUROC 0.906 ± 0.044). A conservative GI-restricted gene product presence/absence proxy achieved AUROC 0.887 ± 0.083, while a full joint-run pangenome GPA baseline remains a necessary future benchmark. Fragmentation-controlled analyses confirmed cassette signal remained informative after quality filtering (AUROC 0.827 in assemblies with ≤50 contigs; n = 91), while leave-one-BioProject-out validation yielded AUROC 0.694, indicating that deployment in novel surveillance contexts requires prospective validation. SHapley Additive exPlanations (SHAP) analysis localized discriminant signal to GI-anchored modules enriched for AMR cargo, mobility load, and GI AMR density.

CONCLUSION: These results suggest that cassette architecture captures signal consistent with biologically meaningful genomic organization beyond bulk island burden and supports its use as an interpretable exploratory representation for surveillance-oriented analysis of poultry-associated E. cecorum, while prospective validation in independent surveillance collections and a full joint-run pangenome gene presence/absence benchmark remain necessary before operational deployment claims can be made.

RevDate: 2026-07-23
CmpDate: 2026-07-23

Timkina E, Palyzová A, Marešová H, et al (2026)

Genomic signatures of radiation stress adaptations in Kocuria rhizophila: insights from strain 301 of the Jáchymov radon springs.

Frontiers in microbiology, 17:1814458.

Several strains of Kocuria rhizophila have been reported to tolerate ionizing radiation and other environmental stresses; however, this phenotype is unevenly distributed across the species. Here, we characterize K. rhizophila strain 301, isolated from chronically radioactive radon springs in Jáchymov (Czech Republic). Strain 301 displayed exceptional stress resilience, retaining approximately 10% viability after exposure to 1.0 kGy of γ-irradiation, and maintaining ~20% survival following 30 days of desiccation. These responses sharply contrasted with the pronounced sensitivity of the type strain K. rhizophila TA68. Complete genome sequencing produced a single circular chromosome of 2.77 Mbp. Comparative genomic analyses revealed extensive duplication of genes involved in DNA repair, antioxidant defense, and metal ion homeostasis, including multiple paralogs of uvrA, uvrD, sodA, and Mn/Fe transport systems. In contrast to the paradigm established by Deinococcus radiodurans, strain 301 lacks radiation-specific or lineage-exclusive genes, instead suggesting that resilience may be associated with quantitative reinforcement of conserved cellular pathways. Pan-genome analysis further demonstrated a closed K. rhizophila pangenome, with strain 301 forming a distinct phylogenetic lineage. Together, these findings position K. rhizophila 301 as a model system of adaptation to chronic radiation exposure and illustrate how sustained environmental pressure may promote modifications and adaptations based on core functions rather than novel innovative genetic traits.

RevDate: 2026-07-20
CmpDate: 2026-07-20

Scherer J, Corá RK, Bonatto D, et al (2026)

Pangenome Dynamics and Functional Diversification in the Marine Genus Pseudoalteromonas: Association to Colony Pigmentation.

Marine biotechnology (New York, N.Y.), 28(4):.

Pseudoalteromonas species are ecologically versatile marine bacteria widely recognized for their capacity to synthesize diverse bioactive metabolites and psychrophilic enzymes with biotechnological relevance. Here, we present a comprehensive comparative genomic analysis of 53 reference genomes to elucidate the dynamics, functional diversity, and biosynthetic potential of this genus. Reference genomes representing each deposited species were selected in order to avoid bias associated with unequal numbers of genomes per species. Pangenome reconstruction revealed an open structure comprising a small core genome (1,350 gene families) and a large proportion of accessory and strain-specific genes, reflecting extensive genomic plasticity. Functional annotation indicated that accessory regions are enriched for genes involved in secondary metabolism, stress adaptation, and environmental resilience. Notably, biosynthetic gene cluster (BGC) mining uncovered a rich repertoire of potentially novel RiPPs and other secondary metabolite clusters, underscoring Pseudoalteromonas as a promising source of unexplored bioactive compounds. Statistical analyses revealed that pigmented strains harbor significantly higher numbers of BGCs compared to non-pigmented strains, while only a weak and non-significant correlation was observed between BGC abundance and carbohydrate-active enzyme (CAZyme) content. No significant effect of the isolation source was detected on either BGC or CAZyme distributions. Together, these findings provide new insights into the genomic basis of ecological adaptation and metabolic diversification in Pseudoalteromonas, supporting the role of pigmentation as a proxy for enhanced biosynthetic potential, while carbohydrate utilization capabilities evolve more independently and offering a framework for targeted bioprospecting of marine-derived metabolites with industrial and environmental applications.

RevDate: 2026-07-20

Tan B, Zafra C, C Ng (2026)

Comparative genomics of the Nap2-2B clade reveals substrate partitioning and niche diversification among uncultured hydrocarbon-degrading Desulfotomaculales.

Scientific reports pii:10.1038/s41598-026-63016-x [Epub ahead of print].

Uncultured Nap2-2B bacteria (order Desulfotomaculales; formerly family Peptococcaceae) are frequently detected in methanogenic hydrocarbon-degrading environments, yet their metabolic diversity remains poorly understood. Here, we analysed 17 GTDB r232 metagenome-assembled genomes (MAGs) from four genera within this clade. A bac120 phylogeny places Nap2-2B as a monophyletic family-level lineage within Desulfotomaculales. Glycyl radical enzyme phylogeny and operon context reveal strict substrate partitioning: SCADC1-2-3 encodes alkylsuccinate synthase for aliphatic hydrocarbon activation, 46-80 and UBA4053 encode benzylsuccinate synthase for aromatic activation, and JAIMBK01 lacks hydrocarbon activation genes but retains complete dissimilatory sulfate reduction pathway genes. Pangenome-level pathway reconstruction identifies complementary cofactor biosynthetic potential, notably in cobalamin and pantothenate biosynthesis, consistent with possible cofactor complementation. Genome-scale metabolic modeling suggests that the alkane-degrading SCADC1-2-3 lineage can support syntrophic hexane degradation, whereas the aromatic lineage cannot grow on the alkane FBA test because it lacks AssA and PFOR. A parallel aromatic-substrate FBA for 46-80 MAGs did not yield growth under minimal curation, reflecting the greater complexity of the downstream benzoyl-CoA pathway. Together, these data support a syntrophic guild structured by substrate partitioning, putative cofactor complementation, and distinct electron-disposal strategies that may shape methanogenic hydrocarbon attenuation in anoxic tailings environments.

RevDate: 2026-07-20

Rathna V, Kukreti A, Prasannakumar MK, et al (2026)

Correction: Pan-genome and antibiotic resistance insights into Xanthomonas citri pv. punicae pathotypes.

BMC microbiology, 26(1):.

RevDate: 2026-07-21

Guan J, Li X, Miao H, et al (2026)

Pangenome-resolved structural variation drives adaptation and trait evolution in cucumber.

Nature genetics [Epub ahead of print].

Cucumber (Cucumis sativus L.) is a global vegetable crop and powerful model for sex determination, fruit development and vascular biology. We present high-quality genome assemblies for 125 cultivated and wild accessions, capturing worldwide genetic diversity. Syntenic gene family analysis characterized 37,897 gene families and revealed haplotype diversity shaped by geographic expansion. Comparative analyses uncovered copy-number variations linked to local adaptation, including a CsFT tandem duplication promoting early flowering at higher latitudes. This resource reduces reference bias, enabling the annotation of resistance loci and the discovery of CsCcu, a nucleotide-binding leucine-rich repeat-type R gene conferring scab resistance. We cataloged 135,597 structural variations and quantified their regulatory effects, with ~30% driving trait diversification among geographic groups. Integrating structural variations into genome-wide association studies identified 172 quantitative trait loci for 38 agronomic traits, including a rare long terminal repeat insertion regulating fruit length via CsSPL1. Our findings provide a genomic toolkit for cucumber evolution research and precision breeding.

RevDate: 2026-07-18

Muscò A, Longhi G, Selleri E, et al (2026)

Mucin O-glycan degradation by GH101 underpins mucosal persistence of Bifidobacterium bifidum PRL2010.

Applied microbiology and biotechnology pii:10.1007/s00253-026-13964-1 [Epub ahead of print].

Host-derived mucin O-glycans constitute a key chemical component of the human intestinal niche and are continuously encountered by gut-adapted bifidobacterial cells, thereby playing a pivotal role in host-microbe interactions. Among these microbes, Bifidobacterium bifidum is one of the most persistent members of the human gut microbiota. Comparative genomic analyses of mucosa-associated strains of this species have revealed a gene conserved across its pangenome, predicted to encode a GH101 glycoside hydrolase involved in mucin degradation. In this study, we performed a molecular characterization of the GH101 enzyme encoded by B. bifidum PRL2010. Insertional mutagenesis of the GH101 gene from the PRL2010 genome resulted in a marked reduction in bacterial adhesion to mucin-secreting epithelial cells, along with a significant growth difference when N-acetyl-galactosamine was provided as the sole carbon source. These findings indicate that GH101 mediates the initial cleavage of O-GalNAc core 1 structures, representing a crucial step for the adhesion to human mucosa. However, this enzyme represents only one component of the broader set of glycosidases required for the complete degradation of host mucins. Overall, our results establish a mechanistic link between metabolic specialization and ecological fitness, providing insights into how mucin-adapted commensals may contribute to interactions with the human gut mucosa. KEY POINTS: • The GH101 glycoside hydrolase of Bifidobacterium bifidum PRL2010 plays a pivotal role in host mucin utilization by initiating the cleavage of O-GalNAc core 1 structures. • Insertional mutant of the GH101 gene impairs both mucosal adhesion and growth on N-acetyl-galactosamine. • These findings reveal a mechanistic connection between mucin glycan metabolism and ecological fitness by mucin-adapted bifidobacterial commensals.

RevDate: 2026-07-18
CmpDate: 2026-07-18

Sisay T, Berhan A, Mihrete K, et al (2026)

Genomic regulation of the diphtheria toxin gene and Its implications for molecular diagnostics and surveillance in low-resource settings.

Molecular biology reports, 53(1):.

Corynebacterium diphtheriae remains a significant, though often underestimated, public health concern, particularly in low- and middle-income countries. The pathogenicity of the disease is primarily determined by diphtheria toxin (DT), which is produced by the tox gene, a bacteriophage-associated element, and is tightly regulated by the iron-dependent transcriptional repressor DtxR, encoded by the dtxR gene. Despite extensive investigation into the molecular biology of DT, its regulation within the broader genomic organization, as well as its implications for diagnostic methods and surveillance strategies, have not yet been fully elucidated. This review consolidates existing evidence regarding the genomic context and molecular regulation of the tox gene, encompassing chromosomal organization, variability in GC content, genomic islands, and mechanisms of horizontal gene transfer. Significant attention is focused on lysogenic conversion mediated by corynephages and regulatory pathways responsive to iron. We also evaluate both established and novel molecular diagnostic approaches, including PCR, real-time PCR, sequencing technologies, and isothermal amplification methods like loop-mediated isothermal amplification (LAMP). Recent genomic discoveries, including pan-genome variation, CRISPR-Cas mechanisms, and the emergence of non-toxigenic tox-bearing strains are analyzed in relation to diagnostic precision and epidemiological surveillance. Understanding the genomic regulation and evolutionary dynamics of toxin production is essential for improving diagnostic accuracy and strengthening surveillance systems, particularly in resource-limited settings where diphtheria is often underdiagnosed and underreported.

RevDate: 2026-07-18

Singh J, Gudi S, Maughan PJ, et al (2026)

A reference-grade chromosome-level genome assembly of a historically important U.S. bread wheat cultivar Timstein.

Scientific data pii:10.1038/s41597-026-07930-9 [Epub ahead of print].

We report a near telomere-to-telomere high quality genome assembly of the historically important spring wheat cultivar, Timstein, generated using PacBio HiFi long-read sequencing data followed by Hi-C scaffolding. The assembly spanned 14.76 Gb, accounting for all 21 chromosomes of the A, B, and D subgenomes. Gene annotations identified around 105 K high-confidence (HC) gene models. The genome was comprised of ~85% transposable elements, primarily from the Gypsy, Copia, and CACTA families. For each subgenome, the BUSCO completeness score ranging from 97.4 to 99.6% and LTR Assembly Index (LAI) values surpassing 13 indicated the assembly quality was reference grade. Synteny analysis with IWGSC Chinese Spring (CS) RefSeq v2.1 revealed strong chromosomal collinearity between two genomes. Timstein has been extensively studied in classical genetics research for stem and leaf rust resistance and septoria nodorum blotch (SNB) susceptibility. This high-quality genome assembly provides cultivar-resolved references that can expand the wheat pangenome, supports structural and functional genomics studies, and enables fine mapping of disease resistance/susceptibility loci for their utilization in wheat genetics and breeding programs.

RevDate: 2026-07-20
CmpDate: 2026-07-20

Young CE, O'Sullivan H, Alattas H, et al (2026)

Refining Salinivibrio pangenome dynamics and biotechnological potential through comparative analysis.

Microbial genomics, 12(7):.

Current understanding of genomic diversity within the halophilic genus Salinivibrio relies predominantly on draft genomes, with only seven complete genomes among the 62 publicly available. Previous pangenome analysis suggested a closed genomic structure while concluding that Salinivibrio lacks polyhydroxyalkanoate (PHA) degradation capacity despite possessing biosynthesis genes. Here, we present eight complete Salinivibrio genomes from Pearse Lakes (Rottnest Island, Western Australia) generated using Oxford Nanopore long-read sequencing, alongside re-analysis of 38 high-quality public genomes (≥90% completeness and ≤5% contamination cut-off). Pangenome analysis revealed a more open structure than previously reported, with a core genome comprising 25% of total gene clusters and an accessory genome accounting for 71%. Panstripe analysis demonstrated significant temporal signal in gene gain and loss events associated with phylogenetic branch length (core: P=1.72×10[-4]; tip: P=2.64×10[-14]). All 46 genomes contained complete PHA biosynthesis operons (phaB-phaA-phaP-phaC) with high sequence conservation under strong purifying selection (Z=30.30, P<0.001). In a genome that readily gains and loses genes, this conservation indicates that PHA synthesis is a maintained pathway, which is difficult to reconcile with a previous report that Salinivibrio lacks PHA degradation capacity. We therefore searched the genomes by Hidden Markov Model-based homology rather than standard annotation and identified seven putative depolymerases that form a single accessory cluster in 15% of strains, all previously annotated as 3-oxoadipate enol-lactonase-2. These candidates retained all catalytic residues characteristic of active depolymerases but are divergent from reference PHA depolymerases which could explain why annotation missed them. They remain putative and require biochemical confirmation. Both the expanded pangenome and these candidates emerged from standardized homology-based re-analysis, showing that annotation-dependent approaches can overlook genomic diversity and divergent enzyme families in non-model organisms. Together, these results establish Salinivibrio as a genomically dynamic genus with potential for halophilic bioplastic production.

RevDate: 2026-07-17
CmpDate: 2026-07-17

Sirén J, Paten B, Human Pangenome Reference Consortium (2026)

GBZ-base and GAF-base: Indexed pangenome file formats.

bioRxiv : the preprint server for biology pii:2026.07.10.737775.

MOTIVATION: Existing pangenome file formats are designed for batch processing. Graphs must be loaded into memory, and alignment files must be read sequentially. Indexed file formats that can be used directly from disk would be more appropriate for interactive applications.

RESULTS: We propose GBZ-base and GAF-base - SQLite-backed file formats comparable to GBZ and GAF. GBZ-base supports efficient extraction of local subgraphs, and GAF-base lets us extract all alignments to the subgraph. Additionally, GAF-base is smaller than any other file format for sequence-to-graph alignments.

From https://github.com/jltsiren/gbz-base and https://crates.io/crates/gbz-base under the MIT license.

RevDate: 2026-07-17
CmpDate: 2026-07-17

Hunt M, Torres MDT, Alikhan NF, et al (2026)

AllTheBacteria: a community resource empowers biology and discovers novel peptide antibiotics.

bioRxiv : the preprint server for biology pii:2024.03.08.584059.

Public microbial genomes encode an immense record of biological diversity, evolution and molecular function, but much of this information remains difficult to reuse because raw sequencing data are not uniformly assembled, quality controlled, annotated or searchable at scale. Here we present AllTheBacteria, an open, community-built resource that transforms public bacterial short-read whole-genome sequencing reads into a uniformly processed discovery platform. The current analysed release contains 2,440,377 high-quality bacterial and archaeal genomes from 11,273 species, together with standardized taxonomic assignments, genome annotations, antimicrobial resistance calls, antiphage-defence annotations, protein structure predictions and AI-ready sequence tables. We show that this infrastructure enables applications that would otherwise be impractical, from global sequence search and outbreak contextualization to pangenome method development, antimicrobial resistance reservoir mapping and antiphage-defence ecology. As a stringent experimental demonstration, we mined 3,919,096 encrypted peptide fragments from AllTheBacteria proteomes using our deep learning model APEX 1.1, identifying 1,867 candidates with predicted antimicrobial activity. We synthesized 24 representative peptides and tested them against 20 clinically relevant bacterial strains, including antibiotic-resistant pathogens. Multiple peptides showed low-micromolar activity, membrane-responsive conformational transitions and selective envelope perturbation. A lead molecule, ATB20, reduced Acinetobacter baumannii burden in a murine skin abscess model with efficacy comparable to polymyxin B and no overt toxicity. Together, these results establish AllTheBacteria as both a foundational community resource for microbiology and a renewable engine for AI-guided antimicrobial discovery.

RevDate: 2026-07-17
CmpDate: 2026-07-17

Bhure M, Shukla N, Purohit H, et al (2026)

Genome-wide investigation of outbreak-associated Vibrio cholerae in Gujarat, India identifies antimicrobial resistance genes, virulence determinants, and mobile genetic elements.

Frontiers in microbiology, 17:1851551.

This study investigates the 2024 cholera outbreak in Gujarat, India, utilizing combined whole-genome analysis of clinical Vibrio cholerae isolates and wastewater surveillance. A total of, 69 V. cholerae isolates were recovered from affected patients, predominantly belonging to the O1 serogroup (51 isolates). Antimicrobial susceptibility test (AST) of 34 isolates revealed complete resistance to ampicillin and partial resistance to cotrimoxazole, whereas all isolates were susceptible to doxycycline, ciprofloxacin, chloramphenicol, tetracycline, and gentamicin. Whole-genome sequencing of 20 selected isolates revealed that the isolates belong to the seventh pandemic El Tor (7PET) lineage, sequence type ST69. Phylogenomic analyses using a multi-method approach, core genes, Composition Vector (CV) Tree, SNPs, and multilocus sequence typing (MLST) showed tight clustering with limited diversity among the isolates. All isolates contained 13-15 antimicrobial resistance genes, with high consistency between genotype-phenotype for most antibiotics, although discordance was observed for ciprofloxacin, cotrimoxazole, and chloramphenicol. Sixteen genes were identified as virulence factors, and 11 isolates also had ctxA/ctxB. All isolates also had two to four integrative conjugative elements (ICEs) containing antimicrobial resistance genes (ARGs) and important Vibrio cholerae pathogenicity islands (VPI-1, VPI-2) and Vibrio cholerae seventh pandemic islands (VSP-1, VSP-2). The pangenome analysis highlights extensive genomic flexibility within species, likely driven by horizontal gene transfer and ecological adaptation; however, further outbreak-specific investigations are required to determine their direct role in current outbreak. The detection of ctxA-positive signals in wastewater, 20% (28/140) of the samples, suggests a possible surveillance signal during the outbreak. These results highlight the presence of antimicrobial-resistant 7PET O1 El Tor strains in Gujarat outbreaks and support continued genomic monitoring to guide focused public health interventions in endemic areas. Furthermore, this study also underscores the importance of wastewater surveillance for monitoring V. cholerae.

RevDate: 2026-07-16

Zhao P, Peng C, Gao Y, et al (2026)

Pangenome Graph Reveals the Structural Variation Landscape in 2929 Cattle Samples and Its Impact on Gene Regulation.

Genomics, proteomics & bioinformatics pii:8735849 [Epub ahead of print].

Structural variations (SVs) represent a significant source of genomic diversity, with demonstrated roles in livestock gene expression and traits. However, a comprehensive understanding of the SV landscape across large sample sets and its impact on gene regulation in cattle remains incomplete. This study aimed to construct high-fidelity pangenome graphs by integrating both assembly-based and whole-genome sequencing (WGS) derived SV catalogs. We evaluated the efficacy of pangenome graphs for SV genotyping and identified 80,328 high-quality SVs from a cohort of 2929 samples. We systematically characterized these SVs, including their linkage disequilibrium with single nucleotide polymorphisms (SNPs), functional annotations, formation mechanisms, and genomic distributions. Furthermore, we generated paired WGS (24.4 ×) and blood RNA-seq data in 170 Simmental cattle. Utilizing our pangenome graphs, we identified 637 SV-expression quantitative trait loci (SV-eQTL), which accounted for 10.81% of expression heritability of target genes, with 38.09% of the effects linked to promoter/enhancer regions. Forty-six of these SV-eQTL were replicated using CattleGTEx results through SV imputation using a joint SNP-SV reference panel. Notably, insertions in the GHSR gene were significantly associated with its expression levels, likely linked to Bos indicus cattle adaptation to heat tolerance. Our findings provide novel insights into the SV landscape and its contribution to gene regulation, underscoring its importance in cattle genetics and genomics.

RevDate: 2026-07-16

Michoud G, Geers A, Peter H, et al (2026)

Evolutionary radiation of Polaromonas from mountain glaciers downstream.

Current biology : CB pii:S0960-9822(26)00816-X [Epub ahead of print].

Habitat transitions are central to microbial ecology and evolution and have been extensively studied across vastly different environments, such as between saline and non-saline environments. However, microbial habitat transitions along other large-scale environmental gradients remain poorly studied. This is particularly true for transitions involving the cryosphere, despite building evidence suggesting the Cryogenian as important for evolutionary radiation. Here, we investigated ecosystem transitions and the related genomic adaptations of the cosmopolitan cryospheric Polaromonas bacterium. We constructed a pangenome from 282 high-quality genomes, sourced from glaciers, glacier-fed streams (GFSs), lakes, wetlands, groundwater, rivers, and soils. Phylogenetic reconciliation suggested that the ancestral Polaromonas genome radiated from glacier ecosystems into various downstream environments through multiple independent transitions. These transitions were likely marked by extensive horizontal gene transfer and gene loss, with mobile genetic elements such as plasmids and prophages playing key roles in genomic diversification. Predicted ancestral genomes encoded versatile metabolic and stress-response capacities, which support adaptation to fluctuating and extreme conditions in the various cryospheric habitats. Compared to the ancestral Polaromonas genome, distinct genomic signatures were associated with specific habitats: GFS lineages possess expanded stress-tolerance repertoires, glacier lineages gained chemolithotrophic and anaerobic pathways, lake and wetland genomes acquired phototrophic functions, and soil lineages expanded substrate transport and stress tolerance. Together, our findings highlight the role of genomic plasticity in the ecological success of Polaromonas and also underscore the cryosphere as a potential evolutionary cradle from which lineages dispersed and adapted to downstream aquatic and terrestrial environments.

RevDate: 2026-07-16

Ashraf H, Doerr D, Ebler J, et al (2026)

Building and applying pangenome references to capture genetic diversity.

Nature reviews. Genetics [Epub ahead of print].

Reference genomes serve as a coordinate system and are central to almost all analyses in genomics. However, linear reference genomes are based on a single individual or a small number of individuals and do not represent genetic diversity. Recent advances in de novo genome assembly, powered by long-read sequencing technologies, now enable the sequence reconstruction of many genomes to reference quality. These pangenomes integrate sequences from multiple individuals into graph-based or multi-haplotype representations, capturing genetic variation beyond a single linear reference. The widespread adoption of such pangenome references, which encode a diverse set of haplotypes, thus removes biases and enables the discovery of variants relative to all included haplotype backgrounds. The emergence of corresponding computational tools for analysing structural variants and complex genetic loci opens up opportunities in genome-wide association studies and rare-disease genetics. Here we review these opportunities, as well as challenges concerning pangenomes that need to be addressed by the research community.

RevDate: 2026-07-17

Wang H, Liang J, Li X, et al (2026)

A pan-genomic perspective: comprehensive dissection of the FIG superfamily unravels its evolution, expansion, and environmental adaptation in wheat.

BMC plant biology pii:10.1186/s12870-026-09530-6 [Epub ahead of print].

BACKGROUND: As a major global food crop, wheat faces dual pressures from population growth and deteriorating agricultural environments in efforts to increase its yield. The members of the FIG superfamily are associated with photosynthetic carbon metabolism and stress-responsive pathways in plants.

RESULTS: In this study, we conducted the first systematic pan-genomic analysis to investigate the evolution and function of the FIG superfamily in wheat. The results revealed that this family underwent significant expansion during plant evolution from aquatic to terrestrial habitats and from lower to higher forms, with functional divergence apparently predating the green algal stage as inferred from phylogenetic patterns. The members of this family were primarily derived from three ancestral species, and their expansion was largely driven by whole-genome duplication. Most members were found to be under purifying selection, whereas the TaVTC4-4B was subjected to positive selection. This gene is constitutively highly expressed in green tissues, and its promoter is enriched with cis-regulatory elements associated with light responsiveness, JA/ABA signaling, and stress responses. Expression analysis indicated its strong responsiveness to salt stresses.

CONCLUSIONS: This study elucidates the evolutionary trajectory and functional landscape of the wheat FIG superfamily, laying a theoretical foundation for the potential genetic improvement of photosynthetic efficiency and stress resilience in wheat.

RevDate: 2026-07-17

Solun GK, Dogrusoz U, Bingöl Z, et al (2026)

PG2: algorithms and a web-based tool for effective layout and visual analysis of pangenome graphs.

BMC bioinformatics pii:10.1186/s12859-026-06555-4 [Epub ahead of print].

BACKGROUND: The advent of cost-effective whole-genome assembly has enabled the creation of comprehensive pangenomes with resolved haplotypes across various organisms. This technological leap drives the refinement of tailored methodologies to manage the intricate sequences and variations in extensive collections of related genomes. These methodologies often utilize graphical representations of pangenomes to enhance algorithms for tasks like sequence alignment, visualization, and functional genomics. Leveraging the insights provided by pangenomes, these approaches exhibit improved efficiency in bioinformatics tasks such as read mapping, variant calling, and genotyping. Pangenome graphs are positioned to become invaluable assets in genomics, offering seamless reconciliation of diverse sequence and coordinate systems. While their potential to replace linear reference genomes is uncertain, their adaptability ensures their utility in future pangenomic models. Utilizing graphs for visual representation aids in exploring critical insights and identifying key patterns, facilitating analysis by highlighting connections, trends, and patterns for improved comprehension.

RESULTS: Towards this goal, we present some algorithms to effectively layout and visually analyze pangenome graphs. Then, we introduce an open-source, flexible, easy-to-use, web-based platform named PG2 (PanGenoGrapher), realizing these algorithms.

CONCLUSIONS: Our primary objective here is to incorporate the capabilities of advanced visualization techniques into the analysis of pangenome graphs, thereby enhancing their utility and accessibility for researchers and practitioners in the field. The source code and user guide are openly available on GitHub at https://github.com/iVis-at-Bilkent/pangenographer. A publicly accessible sample deployment is hosted at http://pg2.cs.bilkent.edu.tr. In addition, a demonstration video illustrating the primary use cases of PG2 is available at https://www.youtube.com/watch?v=yCd7-aGY6CQ.

RevDate: 2026-07-17
CmpDate: 2026-07-17

Rana P, Chaudhary C, Pruthi R, et al (2026)

Molecular insights and translational opportunities to enhance heat tolerance in rice.

The plant genome, 19(3):e70281.

Heat stress is an increasingly serious threat to rice (Oryza sativa L.) productivity, yet the genetic and regulatory architecture underlying thermotolerance remain poorly resolved and fragmented across studies. Earlier research focused on individual pathways or specific developmental stages; however, recent advances now support an integrated understanding of heat stress adaptation in rice. This review synthesizes emerging insights into molecular physiology, regulatory signaling, epigenetic memory, and genome-scale variation associated with thermotolerance. We highlight the interconnected roles of calcium reactive oxygen species signaling, heat shock transcription factor networks, translational regulation, and chromatin-based stress memory in shaping reproductive-stage tolerance and maintaining grain quality under elevated temperatures. The review also emphasizes the value of pangenome analyses and structural variant discovery for identifying heat-responsive genes and regulatory elements absent from single-reference genomes. In addition, genome-wide association studies, haplotype-based breeding, genomic selection, and CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)-based genome editing are discussed as promising approaches for functional validation and deployment of favorable alleles controlling polygenic heat resilience. Despite these advances, several challenges continue to hinder translation into breeding-ready outcomes, including limited field-based validation of candidate genes, poor integration of multi-omics datasets into predictive breeding frameworks, and insufficient understanding of reproductive-stage regulatory networks. Furthermore, genotype × environment interactions, together with trade-offs among yield, grain quality, and stress resilience, strongly influence the stability and transferability of thermotolerance traits across diverse agroecological environments. By integrating mechanistic insights with genome-scale diversity and predictive breeding tools, this review outlines a genomics-enabled roadmap for developing heat-resilient rice cultivars under intensifying global warming and supporting sustainable global rice production.

RevDate: 2026-07-17
CmpDate: 2026-07-17

Guarracino A, Gyamfi A, Human Pangenome Reference Consortium, et al (2026)

Concerted evolution and unorthodox recombination of human subtelomeres.

bioRxiv : the preprint server for biology pii:2026.07.10.737660.

Human subtelomeres contain duplicated sequence that is shared among the ends of non-homologous chromosomes and provides a substrate for ectopic exchange [1-6]. However, incomplete reference assemblies and chromosome-by-chromosome analyses have prevented a population-scale view of the extent and organization of subtelomeric exchange [7-9]. Here we apply a reference-free pangenome approach to 465 near-complete human assemblies, comparing every chromosome end against every other, and find that high-identity pseudo-homolog regions occur on 41 of 48 chromosome arms. These regions form structured sequence communities in which previously described exchange systems appear as local peaks within a broader continuum. Human and mouse chromosome-contact maps show preferential proximity between subtelomeres with similar sequences. In mouse meiosis this proximity is strongest at the zygotene bouquet, when telomeres cluster at the nuclear envelope; in human data it persists even in adjacent flanks that lack the shared sequence used to define each pair. In a three-generation telomere-to-telomere pedigree, whole-genome comparison identifies putative recombination between subtelomeric regions on non-homologous chromosomes that matches this community organization, while recovering the obligate Xp/Yp PAR1 recombination in the male germline. These results generalize known subtelomeric exchange systems into a near-ubiquitous architecture and support recurrent ectopic exchange as a genome-wide force in the concerted evolution of human chromosome ends.

RevDate: 2026-07-17
CmpDate: 2026-07-17

Shivakumar VS, Langmead B, Human Pangenome Reference Consortium (2026)

Navigating the pangenome coordinate system with Shredtools.

bioRxiv : the preprint server for biology pii:2026.07.03.736354.

Existing notions of pangenome coordinates rely on hard-to-compute multiple sequence alignments. On the other hand, pangenome-wide exact unique matches (multi-MUMs) can be computed efficiently, and represent conserved stretches of columns in the underlying MSA. We introduce Shredtools, which uses multi-MUMs as pangenome waypoints and allows for sophisticated queries in pangenome coordinates. Its primary query is extract , which takes an interval of one sequence and extracts the smallest window containing it that is syntenic pangenome-wide. Shredtools' extract query can extract a gene region from 476 human genomes in half a second. Other queries help to refine these results, by finding local exact matches to improve the density of multi-MUM coverage ("enhance") and by selectively discarding sequences to improve the precision of the syntenic region ("zoom"). The Shredtools web interface (available at https://vikshiv.github.io/shredtools) allows for client-side handling of extract queries with index queries handled via simple and fast HTTP Range requests, simplifying usage and enabling pangenomescale discoveries.

RevDate: 2026-07-15
CmpDate: 2026-07-15

Wang G, Zhu N, Sun X, et al (2026)

Pan-Genome and Transcriptome-Guided Analysis Reveals Duplication-Driven Evolution and Candidate MYB-bHLH Modules Associated with Fruit Development in Pear.

Plants (Basel, Switzerland), 15(13): pii:plants15131961.

Gene duplication and subsequent selection are central to genome evolution and transcription factor diversification, but the conservation and divergence of the basic helix-loop-helix (bHLH) family in pear remain unclear from a pan-genome perspective. Here, we performed a pan-genome and transcriptome-guided analysis across 15 pear genome assemblies, including Asian pear, European pear, and hybrid/haplotype assemblies. Genome-wide duplicated gene pairs were classified into different duplication types, and Ka, Ks, and Ka/Ks values were calculated to establish an evolutionary background for duplicated pear genes. Based on this framework, 3222 bHLH were identified and grouped into evolutionary clades and orthologous gene groups. The pear bHLH family contained conserved core members and variable dispensable members, indicating both functional conservation and genome diversification. Duplication and Ka/Ks analyses showed that WGD/segmental duplication contributed to bHLH expansion and that most duplicated PbrbHLH gene pairs were constrained by purifying selection. By integrating 17-tissue and fruit-development transcriptomes from three pear cultivars, 39 fruit-development-associated PbrbHLHs were selected. Co-expression analysis with 185 PbrMYBs identified candidate MYB-bHLH co-expression modules from the available pear fruit-development transcriptomes. These results provide an evolutionary framework for pear bHLH diversification and candidate regulatory modules for future functional studies.

RevDate: 2026-07-15
CmpDate: 2026-07-15

Wu N, Feng Y, Ning X, et al (2026)

Pan-Genome Analysis Reveals Evolutionary Dynamics and Functional Divergence of the NAC Gene Family in Soybean.

Plants (Basel, Switzerland), 15(13): pii:plants15132010.

Soybean (Glycine max) is an important model crop for studying plant functional genes, such as the NAC transcription factor (TF) gene family. The NAC transcription factor (TF) family is one of the largest plant-specific TF families and plays critical roles in plant growth, development, and stress responses. In this study, we performed a pan-genome-wide analysis of NAC genes using 29 soybean genomes. A total of 5051 NAC genes were identified and clustered into 245 orthologous gene groups (OGGs), including 58 core, 88 soft-core, 32 shell, and 67 cloud groups. Based on phylogenetic relationships, the representative NAC OGGs were assigned to 18 subfamilies, 17 of which contained soybean NAC genes. Gene duplication analysis indicated that whole-genome duplication (WGD)/segmental duplication was the predominant driver of NAC family expansion, accounting for 90.88% of duplication events. Approximately 39.30% of NAC genes carried at least one intact transposable element (TE) within 2 kb upstream or downstream regions. NAC genes with copy number variation (CNV) harbored more nearby TEs than non-CNV genes (1.54 vs. 1.31 TEs per gene), and dispensable NAC genes contained more nearby TEs than core NAC genes (1.59 vs. 1.33 TEs per gene). These results indicate a significant association between local TE abundance and NAC gene CNV or dispensability. Selection pressure analysis showed that dispensable NAC genes had higher Ka, Ks, and Ka/Ks values than core genes, suggesting relatively relaxed evolutionary constraints. Expression profiling across six tissues revealed distinct transcriptional patterns among NAC subfamilies. Structurally conserved subfamilies generally showed broader expression, whereas structurally divergent subfamilies displayed greater expression variability. Regulatory network and Gene Ontology (GO) enrichment analyses suggested that conserved subfamilies were mainly associated with stress responses, while divergent subfamilies were related to cell wall regulation, signal transduction, and ion homeostasis. Further analysis of Wm82 drought RNA-seq data prioritized several putative drought-responsive NAC candidates, including Glyma.16G043200, Glyma.06G248900, Glyma.07G050600, Glyma.12G206900, and Glyma.18G261300. Overall, these findings elucidate the mechanisms of expansion and the functional divergence of the NAC gene family at the soybean pan-genome level, providing a theoretical basis for understanding NAC gene evolution and facilitating future crop improvement.

RevDate: 2026-07-15

Calvigioni M, Rossi V, Celandroni F, et al (2026)

Identification and in-depth characterization of clinical isolates of Peribacillus frigoritolerans.

Microbiology spectrum [Epub ahead of print].

UNLABELLED: Peribacillus frigoritolerans is a bacterial species commonly found in the environment and used as a plant-growth promoter and biocontrol agent in agriculture. Recent evidence has proven that Peribacillus spp. are also able to cause severe infections in humans, thus emerging as new human pathogens. In this study, for the first time, 10 P. frigoritolerans strains were isolated from human samples (both superficial and sterile deep body sites) and characterized in terms of morphology, lifestyle, genetics, and virulence. The molecular identification by MALDI-TOF mass spectrometry and 16S rRNA gene sequencing was inconclusive, while whole-genome sequencing was effective in properly identifying isolates within the species P. frigoritolerans. The pangenome analysis provided an overview of the virulence potential of P. frigoritolerans, revealing the presence of genes involved in antibiotic resistance and toxin/exoenzyme production. Phenotypically, the strains displayed different features and behaviors, indicating strain-specific properties and high intra-species variability. A part of the strains exhibited virulence factors, being able to swim and swarm, form biofilms, and produce enzymes and toxins. Antibiotic susceptibility testing revealed resistance to ampicillin for all strains and resistance to erythromycin and clindamycin for some of them. Antimicrobial activity against Gram-positive bacteria and fungi was demonstrated, further corroborating the presence of putative bacteriocin/antimicrobial peptide-encoding genes. An association between the overall virulence potential and infection site/severity was hypothesized. Altogether, these findings highlight the extreme diversity within the species, reveal the strain-dependent pathogenic potential of P. frigoritolerans, and support its role as a candidate human pathogen.

IMPORTANCE: This study provides insights into the infectious role of Peribacillus frigoritolerans, an almost unknown bacterial species with agrobiotechnological potential but no history of human infections. This is the first report of P. frigoritolerans isolation from human clinical samples. Ten P. frigorit-olerans strains were herein characterized for their morphology, lifestyle, genetics, and virulence, highlighting an extreme intra-species variability and the potential to act as pathogens in humans. Importantly, this study points out the need for unconventional methods for proper identification of this species, since traditional techniques result inconclusive. Resistance to commonly prescribed antibiotics was also evidenced, confirming the importance of antimicrobial testing on clinical iso-lates. This study lays the foundation for a more in-depth characterization of Peribacillus spp. in the clinical context.

RevDate: 2026-07-16
CmpDate: 2026-07-16

Cárdenas JP, Vidal-Veuthey B, Meza K, et al (2026)

In-silico analysis of Bifidobacterium bifidum strain 900791 genome in the context of the B. bifidum pangenome.

Frontiers in cellular and infection microbiology, 16:1744409.

INTRODUCTION: Bifidobacterium bifidum is a key member of the human gut microbiota with well-recognised roles in intestinal homeostasis, glycan metabolism, and immunomodulation. Strain 900791, isolated from the meconium of a Siberian infant, has been used as a commercial probiotic ingredient for decades and has demonstrated, in previously published clinical trials, an ability to improve lactose tolerance and reduce gastrointestinal symptoms in both children and adults; however, the genomic basis for these properties has not been characterised.

METHODS: We present the complete genome sequence and comprehensive in silico genomic and pangenomic analysis of B. bifidum strain 900791. Hybrid sequencing was used for genome assembly. Phylogenomic analysis, including core genome MLST (cgMLST), integrated 229 high-quality B. bifidum genomes, representing the largest dataset for this species to date. Functional annotation and carbohydrate-active enzyme (CAZyme) profiling, antimicrobial resistance prediction, and bioinformatic screening for probiotic-associated genomic features were also performed.

RESULTS: Hybrid sequencing yielded a single circular chromosome of 2,280,092 bp, comprising 1,852 protein-coding sequences. Phylogenomic analysis revealed that strain 900791 belongs to a clonal subgroup of nine closely related strains (>99% cgMLST identity), consistent with a geographically structured lineage. The species pangenome comprised 4,152 orthogroups and a core of 1,450 gene families; 23 orthogroups were exclusive to the 900791 clonal subgroup, including a predicted lantibiotic biosynthetic cluster. CAZyme profiling identified glycoside hydrolase families associated with human milk oligosaccharide degradation (GH2, GH20, GH33, GH84), mucin glycan cleavage (including ten GH families), and lactose metabolism (GH2, GH42). Safety assessment identified only species-typical resistance to mupirocin and rifampicin, with no acquired resistance markers. Bioinformatic screening of the clonal subgroup detected the presence of adhesion-associated proteins, acid resistance systems, bile salt tolerance determinants, oxidative stress response proteins, and two putative bacteriocin gene clusters.

DISCUSSION: These findings provide a genomic framework consistent with the documented clinical role of strain 900791 in lactose tolerance and support its further investigation as a candidate probiotic. The probiotic-associated features identified here may help explain its observed properties and represent priority targets for experimental validation in future in vitro and in vivo studies.

RevDate: 2026-07-14

Oles RE, Carrillo Terrrazas M, Loomis LR, et al (2026)

Comparative genomic analysis of Bacteroides fragilis from intestinal and extra-intestinal sites.

Microbiology spectrum [Epub ahead of print].

Bacteroides fragilis, a key member of the human gut microbiota, contributes to host health by maintaining intestinal homeostasis. Yet, it is also the most frequently isolated anaerobe in clinical infections. These contrasting roles raise questions about the genetic and ecological factors that explain why this common symbiont is disproportionately linked to infection. We analyzed 813 Division I B. fragilis genomes, including 147 new isolates from intestinal and extra-intestinal sites. Infection-associated isolates spanned all phylogroups, indicating no pathogenic lineage. We identified 16 phylogroups, distinguished by genes associated with capsule biosynthesis and interbacterial competition. Additionally, differential metabolomic analysis identified 12 metabolites associated with isolation source, while a microbial genome-wide association study uncovered 44 genes enriched in isolates from extra-intestinal sites, providing the first population-scale markers tied to clinical recovery sites. These results do not implicate a pathogenic lineage; instead, they point to associational links between accessory modules and recovery from extra-intestinal sites under permissive host conditions. This work underscores how genomic diversity and ecological context may jointly shape the clinical impact of gut commensals.IMPORTANCEBacteroides fragilis, a human gut resident, is paradoxically one of the most frequent anaerobes recovered from bloodstream and abscess infections. The genetic features that enable frequent recovery from extra-intestinal sites remain poorly defined. Using comparative genomic and metabolomic analyses of strains from intestinal and extra-intestinal sources, we show that strains isolated from infections are phylogenetically dispersed rather than restricted to a single lineage. We observe lineage-linked differences in capsule loci and competition systems, which suggests constrained gene flow and lineage-specific adaptation within the gut. Additionally, a subset of genes and metabolites is enriched among extra-intestinal isolates. Together, these findings suggest that extra-intestinal survival among B. fragilis strains reflects the interplay between species-wide genomic diversity and permissive host conditions, rather than the emergence of a single pathogenic lineage.

RevDate: 2026-07-14

Frolova M, Maguire B, Duessmann H, et al (2026)

HumanFilt: a multi-reference host depletion pipeline improves Fusobacterium detection accuracy in tumor WGS data sets.

mSystems [Epub ahead of print].

UNLABELLED: The study of tumor-associated microbiomes using whole-genome sequencing (WGS) has attracted considerable attention, but microbial signal detection remains controversial due to host contamination and methodological artifacts. As the necessity of human-read removal becomes increasingly evident, many groups now include this step in their data pre-processing workflows. In this work, we introduce an open-source tool, HumanFilt, designed for rigorous host-read removal and apply it to the re-analysis of WGS data from 10 mucinous rectal adenocarcinoma cases originally published by Reynolds et al. The workflow integrates k-mer-based classification (Kraken2), quality and adapter trimming (Trim Galore), vector filtering (BBDuk/UniVec_Core), and duplicate removal (FastUniq). After reducing data complexity, a multi-aligner, multi-reference approach (BWA-MEM/GRCh38, Bowtie2/T2T-CHM13, and Minimap2/Human Pangenome Reference Consortium v1.1) removes remaining host sequences, collectively eliminating more than 99.9% of human-derived reads. Although the additional alignment steps eliminated only a small fraction of total reads, they consistently removed millions of residual sequences per sample, underscoring the importance of rigorous filtering in data sets where non-human reads are a small minority. Taxonomic profiling with PathSeq and MetaPhlAn revealed reproducible enrichment of Fusobacterium species in tumor versus matched normal tissues, and comparison before and after filtering showed that this tumor-over-normal pattern was preserved despite an overall reduction in RPM values. Simulation analyses further showed that HumanFilt preserved more than 99.7% of true Fusobacterium signals, supporting high specificity without meaningful false-negative loss of microbial reads. In direct comparison with Deacon and NoHuman, HumanFilt achieved the most stringent host-read removal but also removed a greater proportion of PathSeq-classified Fusobacterium reads, highlighting the trade-off between maximal host depletion and preservation of ambiguous microbial signal. Cross-validation with immunofluorescence analysis using pan-Fusobacterium (detecting both Fusobacterium animalis and Fusobacterium nucleatum) and F. nucleatum-specific antibodies showed general consistency with Fusobacterium subspecies detected by WGS. Compared to the unfiltered analysis, host depletion markedly reduced artificial microbial signals in normal samples while preserving tumor-associated Fusobacterium, resulting in a more reliable microbial profile.

IMPORTANCE: We developed an open-source tool that enables rapid removal of human-derived sequences and applied it to rectal cancer whole-genome sequencing data. This approach reduced false microbial signals while preserving true tumor-associated Fusobacterium, and simulation analyses showed that it retained more than 99.7% of true Fusobacterium reads. Comparison with Deacon and NoHuman showed that HumanFilt achieved more stringent host depletion but also highlighted the trade-off between aggressive host filtering and preservation of ambiguous microbial signal. We also observed general consistency between the sequencing results and immunofluorescence staining in tissue. Together, these findings provide a more reliable basis for studying tumor-bacteria interactions.

RevDate: 2026-07-14
CmpDate: 2026-07-14

Gregory JB, Harrison JW, Uehling JK, et al (2026)

Phylogenetically diverse Mucorales-Mycetohabitans endosymbiotic interactions identified from whole-genome sequencing using a targeted metagenomic assembly pipeline.

Microbial genomics, 12(7):.

Endosymbiotic bacteria of the genus Mycetohabitans are obligate intracellular associates of Mucorales fungi, yet the understanding of their diversity, distribution and evolutionary dynamics is in its infancy. By screening 1,696 public sequencing datasets from Mucorales fungi, we detected Mycetohabitans in 46 fungal accessions spanning 5 host taxa across the fungal genera Rhizopus and Apophysomyces. These included 13 previously unreported associations. Genome reconstruction yielded 38 Mycetohabitans metagenome-assembled genomes (MAGs), of which 34 were of high quality. Incorporating these MAGs into genome-based species delimitation expanded known Mycetohabitans diversity from four to nine species-level clusters, including novel host-associated lineages. Re-examination of fungal host identities revealed frequent misidentification of isolates in fungal collection catalogues and/or misannotation in GenBank, with nearly a quarter of positive datasets requiring correction through internal transcribed spacer and genome-scale verification. Host-symbiont associations were non-random under this revised framework, with significant structure detected by contingency analysis and ParaFit. MAG-focused pangenome analysis revealed an open pangenome and mosaic lineage-associated functional traits, including variation in metabolism, secretion, cell-envelope systems, metal resistance, antimicrobial-resistance-associated functions and mobile elements. The most distinctive lineage comprised two Apophysomyces-associated MAGs, provisionally named M. apophysomyceticola, which showed pronounced genome reduction compared with other sampled Mycetohabitans spp. and loss of multiple central metabolic, nutrient assimilation, cofactor biosynthesis, catabolic, stress-response and defence pathways, consistent with reduced metabolic flexibility and increased host dependence. Together, these results show that Mycetohabitans symbioses are more geographically widespread, taxonomically diverse and functionally differentiated than previously recognized. More broadly, this work demonstrates the value of public sequencing repositories for uncovering hidden fungal-bacterial symbioses, while emphasizing that repository-derived patterns must be interpreted considering host misidentification, uneven sampling and incomplete metadata. Overall, our work establishes a global framework for Mycetohabitans diversity and function, with implications for fungal ecology, evolution and clinical mycology.

RevDate: 2026-07-15
CmpDate: 2026-07-15

Belkova N, Smurova N, Zugeeva R, et al (2026)

Pathogenic Potential of Pseudoxanthomonas kaohsiungensis Strain IMB-1 Based on Whole-Genome Sequencing.

Biology, 15(13): pii:biology15131010.

Mass spectrometry and high-throughput sequencing have been introduced into clinical bacteriology. We characterized strain IMB-1, previously isolated from the cerebrospinal fluid of a child, as Pseudoxanthomonas kaohsiungensis and analyzed its biological properties, resistance phenotype, and complete genome. The IMB-1 strain displayed amylolytic, weak lipolytic activities, and it exhibited a phenotypic resistance profile only for aminoglycosides. The dDDH calculation based on the complete genome sequence showed that strain IMB-1 was closely grouped with the type strain P. kaohsiungensis DSM 17583, and the dDDH (d4) value was 70.1%. A comparative pan-genome analysis was performed for four P. kaohsiungensis genomes, revealing a substantial shared core genome. The IMB-1 genome contained 508 unique gene clusters, representing the largest strain-specific gene set among the analyzed genomes, suggesting genomic plasticity and adaptation to the host-associated environment. Genome annotation revealed genes responsible for antibiotic, disinfecting agent, and antiseptic resistance. Gene clusters exhibiting the potential to form biofilms, adhere to the epithelial surface, and exhibit resistance to stress factors were identified. Our study demonstrates that strain IMB-1 is a potential opportunistic pathogen with significant pathogenic potential. The application of high-resolution whole-genome sequencing data in public health for pathogen identification and monitoring can improve the accuracy of infection source determination, reduce the scale and burden of outbreaks, and identify and quantify antimicrobial resistance in pathogens.

RevDate: 2026-07-13

Chaichana N, Singkhamanan K, Wonglapsuwan M, et al (2026)

Genome-guided characterization of Weissella species highlights strain-specific functional traits relevant to food fermentation and probiotics.

Microbiology spectrum [Epub ahead of print].

UNLABELLED: The genus Weissella comprises heterofermentative lactic acid bacteria (LAB) widely distributed in fermented foods, plant-associated environments, and host-associated niches. However, the functional diversity, safety-related genomic features, and evolutionary dynamics of the genus remain incompletely understood at the genomic level. This study performed a comprehensive comparative genomic analysis of 347 high-quality Weissella genomes representing 16 assigned species to elucidate their genomic diversity, pan-genome architecture, and functional potential relevant to food fermentation, probiotics, and biotechnology. Pan-genome analysis revealed an open pan-genome following a power-law model with an exponent of 0.304, dominated by shell and cloud genes, indicating substantial genomic plasticity and ongoing gene acquisition. In silico safety screening showed that the vast majority of genomes lacked detectable antimicrobial resistance or virulence-associated genes under the applied criteria, indicating a low detectable burden of these genetic determinants across the data set. Moreover, functional annotation revealed conserved carbohydrate-active enzyme (CAZyme) repertoires dominated by glycoside hydrolases (GHs) and glycosyltransferases (GTs), alongside species-specific variation in carbohydrate-binding modules (CBMs). Comparative synteny analysis identified six distinct architectures of exopolysaccharide (EPS) biosynthesis loci, with W. cibaria more frequently harboring regulator-rich, putatively complete EPS gene clusters. Putative probiotic-associated genetic markers were unevenly distributed across species, mainly involving stress resistance, vitamin biosynthesis, GIT tolerance, and oxidative stress resistance, whereas GABA production and antimicrobial metabolite-associated markers were not detected in the data set. Hence, this genome-resolved analysis provides a systematic framework for understanding Weissella diversity and supports a genome-guided, strain-level approach for prioritizing candidate strains for future experimental validation in food, health, and biotechnological applications.

IMPORTANCE: Microbial traits relevant to food fermentation, probiotic performance, and safety emerge from complex interactions among core metabolism, accessory genes, and genome plasticity; however, these relationships remain poorly resolved at the genus scale. By analyzing 347 high-quality genomes spanning 16 Weissella species, this study reveals how an open pan-genome, extensive accessory gene diversity, and lineage-specific gene architectures shape functional potential across this lactic acid bacterial genus. Key traits, including EPS biosynthesis, carbohydrate utilization capacity, stress resilience, and biosynthetic gene cluster distribution, are unevenly structured across species and strains, emphasizing the need for genome-guided strain selection rather than taxonomic assumptions. Integrating phylogenomics, gene content variation, and functional profiling links genome evolution with applied phenotypes. These findings advance understanding of how genomic diversity underpins ecological adaptation and biotechnological performance and provide a blueprint for selecting safe and functionally optimized strains in food and health applications.

RevDate: 2026-07-13

Zeng Z, Wang J, Norbu N, et al (2026)

Comparative genomics clarifies phylogenetic relationships and genome evolution in Hippophae from the Qinghai-Tibet plateau.

Molecular phylogenetics and evolution pii:S1055-7903(26)00160-0 [Epub ahead of print].

The uplift of the Qinghai-Tibet Plateau (QTP) and associated climatic oscillations have shaped plant evolution, yet how adaptive strategies diversify along elevational, moisture, and latitudinal gradients within a single clade remains poorly understood. Hippophae (Elaeagnaceae) is distributed along these gradients, with all members sharing Frankia-mediated nitrogen-fixing symbiosis and dioecy. Here, we assembled chromosome-level genomes of H. neurocarpa, H. rhamnoides subsp. yunnanensis, and H. rhamnoides subsp. turkestanica, integrated them with four published assemblies, covering seven taxa (species and subspecies) of the genus. Whole-genome phylogenies dated the major intra-generic divergences to 6.71-2.83 Ma, coinciding with QTP uplift and Asian monsoon intensification during the late Miocene-Pliocene. Two ancient whole-genome duplications (ca. 35-40 and 25-30 Ma) predated the radiation, with retained paralogs significantly enriched in plant hormone signaling, ABA/MAPK cascades, cold-stress response, and reactive oxygen species metabolism. LTR retrotransposons showed a cross-species insertion peak at ∼ 0.2 Ma, and their flanking genes were repeatedly associated with ABA signaling, cold and UV-B responses, and flavonoid metabolism, suggesting a possible link to Pleistocene oscillations. Pan-genome analysis revealed core gene families comprising ∼ 55% of the pan-genome, while variable families differed along ecological gradients, with high-elevation H. tibetana harboring the highest proportion of species-specific genes. Lineage-specific positively selected genes were enriched in DNA damage repair, translational fidelity, and stress signaling. Together, these findings suggest that multidirectional ecological divergence in Hippophae was associated with ancient duplicate retention, transposon-associated regulatory variation, and lineage-specific selection, exemplifying rapid adaptive diversification in plants of the QTP and adjacent regions.

RevDate: 2026-07-13

Kawarizadeh A, Marenda MS, Bailey KE, et al (2026)

Extensively drug-resistant and extended spectrum β-lactamase-producing Raoultella terrigena from the reproductive tract of a mare in Australia.

Scientific reports pii:10.1038/s41598-026-60203-8 [Epub ahead of print].

Raoultella terrigena is a bacterium commonly isolated from environmental sources, with rare reports of isolation from clinical samples. Multidrug resistance is also rare. In this study, R. terrigena was obtained from the uterine lavage of a thoroughbred mare. Broth microdilution and disc diffusion methods were used to determine susceptibility to antimicrobials. Whole genome sequencing was performed to identify and characterise the presence of mobile genetic elements, antimicrobial resistance and virulence genes. Pan-genome analysis allowed comparison of the genome of this isolate with other R. terrigena isolates. Phenotypic resistance was detected to cefazolin, ceftiofur, cefotaxime, ticarcillin/clavulanic acid, tetracycline, doxycycline, gentamicin, sulfamethoxazole/trimethoprim and chloramphenicol. The genome included two plasmids (IncQ1/IncU and IncFII/rep_cluster_2078 types). Also identified in the IncQ1/IncU plasmid were genes conferring resistance to aminoglycosides, fluoroquinolones, rifamycin, phenicols, cephalosporins, diaminopyrimidines, sulfonamides and tetracycline. Except for enrofloxacin susceptibility, the presence of resistance genes was consistent with observed phenotypes. Genes conferring resistance to antiseptics, biocides and metals were also detected. Phylogenetic analysis of the combined dataset, including the isolate in this study and publicly available genomes, revealed two distinct clades, with equine isolates clustering together with other animal-derived isolates. The emergence of antimicrobial resistance in rare opportunistic pathogens underlines the importance of continuous monitoring of such bacteria and emphasizing the need for developing antimicrobial stewardship programs in veterinary settings.

RevDate: 2026-07-13

Chen H, Xing L, Guan C, et al (2026)

Graph-based pan-genome reveals structural variations associated with agronomic traits in mung bean.

Nature genetics, 58(7):1696-1710.

Mung bean (Vigna radiata) is a globally important legume crop valued for its short growing cycle, nitrogen-fixing capacity and high nutritional value, particularly in developing countries. Here we report a comprehensive graph-based pan-genome assembled from 11 genetically diverse global accessions. The framework captures 75,268 gene families (50.86% core, 35.19% dispensable and 13.95% private) and 66,862 nonredundant structural variants. Integrating these structural variants and single nucleotide polymorphisms, genome-wide association studies across five environments identified candidate genes for 20 agronomic traits, underscoring the pivotal roles of these variants in driving mung bean domestication and improvement. Mechanistically, we demonstrate that a 68-bp promoter insertion in VrTIFY6B and a 136-bp promoter deletion in VrPGIP1 regulate flavonoid content and confer bruchid resistance, respectively. These genomic resources and actionable functional variants provide a powerful toolkit to accelerate mung bean improvement through marker-assisted breeding, genomic selection and genome editing to address global food security.

RevDate: 2026-07-11

Adjei MO, Guan C, Dilshad A, et al (2026)

Pan-genomic analysis reveals ecological adaptation and biocontrol of Bacillus pumilus.

Journal of plant physiology, 325:154834 pii:S0176-1617(26)00147-1 [Epub ahead of print].

Bacillus pumilus is known for its ecological resilience and plant-beneficial traits; however, the genomic basis of its biocontrol potential remains unclear. Here, we performed a comparative pan-genome analysis of ecologically diverse B. pumilus strains to explore the genetic determinants underlying plant association and antimicrobial activity. The species exhibited an open pan-genome containing 6035 gene clusters, including a conserved core genome of 3078 clusters and a diverse accessory genome. Functional annotation revealed conserved gene clusters involved in biofilm formation, sporulation, auxin biosynthesis, metal detoxification, short-chain fatty acid regulation, and nutrient-sensing regulators. Genome mining further identified conserved biosynthetic gene clusters encoding antimicrobial compounds and siderophores, including bacilysin and bacillibactin, which are associated with pathogen suppression and plant protection. Several of these strains also possessed unique gene clusters linked to nutrient acquisition, metal detoxification, and environmental adaptation. The universal presence of the chloramphenicol resistance gene (cat86) underscores a conserved adaptive trait. These findings indicate that the ecological versatility of B. pumilus is driven by a stable core genome combined with a dynamic accessory genome enriched in its secondary metabolite pathways, highlighting its potential for sustainable agriculture and biological control. This study presents original research findings based on comparative pan-genome analysis of Bacillus pumilus strains.

RevDate: 2026-07-11

Corbín-Agustí P, Álvarez-Herrera M, Román-Écija M, et al (2026)

A metabolic model based on a pangenome core reveals putative conserved biochemical features of the phytopathogen Xylella fastidiosa.

Microbiological research, 312:128616 pii:S0944-5013(26)00180-1 [Epub ahead of print].

Xylella fastidiosa is a xylem-limited phytopathogenic bacterium responsible for severe diseases in many economically important crops. Despite its impact, its metabolism remains poorly characterized due to fastidious growth and the limited availability of defined culture media. Here, we reconstruct the first pangenome-based genome-scale metabolic model for X. fastidiosa, integrating conserved metabolic functions from 18 strains across five subspecies. The resulting consensus model, iXfcore, is manually curated and used to explore the species' metabolic capabilities. Model simulations predict minimal nutritional requirements that guide us in the formulation of defined media to assess biofilm formation in vitro, supporting the utility of the resulting predictions. Network analysis also identifies a previously undescribed model-predicted candidate pathway for acetate assimilation, consistent with genomic evidence but requiring further empirical validation. In addition, the model predicts the overproduction of polyamines, compounds linked to virulence in other phytopathogens. Experimental analyses confirm polyamine production in multiple X. fastidiosa strains in vitro, providing the first evidence of polyamine detection in culture supernatants of this phytopathogen. Overall, iXfcore provides a systems-level framework to investigate X. fastidiosa metabolism, generate testable hypotheses on its physiology and putative virulence-associated traits, and support future strain-specific models and studies of host-pathogen metabolic interactions.

RevDate: 2026-07-13

Li D, Wang YQ, Huang Y, et al (2026)

A Minimalist Core and Lineage-Specific Expansion Underpin the Adaptive Evolution of the Cyclic Nucleotide-gated Channel (CNGC) Family: Insights from the Tea (Camellia sinensis) Pan-Genome.

Journal of agricultural and food chemistry [Epub ahead of print].

Cyclic nucleotide-gated channels (CNGCs) are critical Ca[2+]-permeable cation channels that orchestrate plant stress responses, yet their evolutionary dynamics in tea plants (Camellia sinensis) remain unknown. Here we dissect the CNGC family using a pan-genome of 28 tea accessions. We identified 867 CNGC genes into 30 orthogroups, revealing a dual evolutionary strategy: an ultraconstrained, two-gene core (CsCNGC8/17) under strong purifying selection ensuring signaling fidelity, and Group IVc─a novel clade (39.4% of the family)─exhibiting lineage-specific motif reconfiguration, dispersed duplications and elevated Ka/Ks ratios. Expression profiling suggested that Group IVc members are broadly upregulated across multiple stresses, whereas core genes display condition-specific responses. This unique architecture─balancing a minimalist, conserved core with a massively expanded, plastic periphery─provides a robust adaptive solution to the lifelong stress exposure characteristic of perennial tea plants. These findings offer critical insights into stress-response evolution and valuable molecular targets for breeding resilient tea cultivars.

RevDate: 2026-07-10
CmpDate: 2026-07-10

Plender EG, Prodanov T, Lin J, et al (2026)

Complex structural variation, phylogeny, and disease associations of the mucin pangenome.

medRxiv : the preprint server for health sciences pii:2026.07.01.26356476.

Mucins are large glycoproteins that provide hydration and barrier function to epithelial tissues. Although genetically heterogeneous, all mucins harbor a large exon composed of variable number tandem repeats (VNTRs). Short-read sequencing has limited our understanding of mucin VNTR diversity and makes disease association studies challenging. We leverage 296 long-read phased genome assemblies to characterize 14 mucin family members, achieving ≥97% accuracy across 572 haplotypes. Phylogenetic haplogroup analysis reveals extraordinary structural heterozygosity, with MUC4 harboring the greatest allelic diversity (n=240 distinct lengths) and MUC12 the greatest size range (Δ = 55,233 bp; 23,080 amino acids). Ten mucins show significant population stratification (pFDR < 0.05). At the MUC4/MUC20 locus, we characterize higher-order structural variation, including a recurrent inversion, copy number variation, and interlocus gene conversion. Optimized genotyping achieves ≥95% haplogroup concordance across 10 loci. We apply this to 4,637 deeply phenotyped cystic fibrosis patients and identify a significant association between short MUC1 VNTRs and severe disease (p=0.0056), demonstrating the pangenome's utility for complex locus genotyping and disease discovery.

RevDate: 2026-07-10
CmpDate: 2026-07-10

Zhu W, Xu G, Gu L, et al (2026)

Phenotypic and comparative genomic characterization of a human biliary-derived Kosakonia radicincitans isolate.

Frontiers in microbiology, 17:1885996.

INTRODUCTION: Kosakonia radicincitans is primarily recognized as a plant-associated and environmentally adapted member of Enterobacteriaceae, whereas recovery from human clinical specimens remains uncommon. Routine biochemical identification may misassign recently reclassified Kosakonia species to closely related Enterobacter taxa.

METHODS: We characterized strain ZJG61129, a K. radicincitans isolate recovered from bile during endoscopic retrograde cholangiopancreatography (ERCP) in a patient with choledocholithiasis. Colony morphology, VITEK 2 Compact identification, MALDI-TOF MS identification and antimicrobial susceptibility testing were performed. Whole-genome sequencing, average nucleotide identity (ANI), digital DNA-DNA hybridization (dDDH), 16S rRNA and core-genome phylogeny analyses, whole-genome comparison, pan-genome analysis, and resistance/virulence-associated gene screening were used for taxonomic confirmation and genomic characterization.

RESULTS: ZJG61129 formed smooth, moist colonies on blood agar and MacConkey agar. VITEK 2 Compact assigned the isolate to the Enterobacter cloacae complex with 92% confidence, whereas MALDI-TOF MS identified it as K. radicincitans with a score of 2.259. The genome consisted of a single circular chromosome of 5,510,622 bp with a GC content of approximately 54.0%. ANI and dDDH analyses confirmed assignment to K. radicincitans, and 16S rRNA together with core-genome phylogeny placed ZJG61129 within the Kosakonia lineage. Comparative genomics showed a conserved genomic backbone, substantial genomic plasticity, and no clear source-associated genomic structuring in the current dataset. ZJG61129 was susceptible to all antimicrobial agents tested. In silico resistance and virulence gene analyses did not indicate high-risk acquired resistance or a clearly high-virulence genotype.

DISCUSSION: ZJG61129 represents a human biliary-derived, environmental-like K. radicincitans isolate that may be misassigned by routine biochemical identification. Genome-based analysis is valuable for accurate recognition of uncommon, recently reclassified Enterobacteriaceae from biliary specimens.

RevDate: 2026-07-10

Arnoux J, Mainguy J, Bry L, et al (2026)

Panorama: A robust pangenome-based method for predicting and comparing biological systems across species.

PLoS computational biology, 22(7):e1013856 pii:PCOMPBIOL-D-25-02697 [Epub ahead of print].

Over the last decade, the expansion in the number of available genomes has profoundly transformed the study of genetic diversity, evolution, and ecological adaptation in prokaryotes. However, traditional bioinformatic approaches based on the analysis of individual genomes are showing their limitations when faced with the sheer scale of the data. To overcome these constraints, the concept of pangenome has emerged, offering a comprehensive framework to capture the full genetic repertoire of a species. In this study, we present PANORAMA, an innovative pangenomic tool designed to exploit pangenome graphs, enabling their annotation and comparison to explore the genomic diversity of several species. Based on the PPanGGOLiN pangenome graphs, PANORAMA integrates advanced methods for rule-based prediction of macromolecular systems and comparative analysis of conserved features between different pangenomes, such as spots of insertion. We illustrate the use of PANORAMA on a dataset of 941 Pseudomonas aeruginosa genomes, evaluating its performance against reference defense system prediction tools such as PADLOC and DefenseFinder. The analysis was then extended to a larger set, including four species of Enterobacteriaceae (>6,000 genomes), demonstrating PANORAMA's ability to annotate, compare, and explore the diversity and distribution of biological systems across multiple species. This work provides new methods for the large-scale comparative study of microbial genomes and highlights the relevance of pangenome approaches in deciphering their evolutionary dynamics. PANORAMA is freely available and accessible at: https://github.com/labgem/PANORAMA.

RevDate: 2026-07-10

Ni Z, Zhang Z, Ning M, et al (2026)

An Anas pangenome graph reveals the role of structural variations in duck domestication.

Poultry science, 105(10):107388 pii:S0032-5791(26)01019-9 [Epub ahead of print].

Despite the significant phenotypic divergence between domestic ducks and their wild progenitors, the structural variants (SVs) underlying this differentiation remain largely unexplored due to the limitations of linear references in capturing genetic diversity across the genus Anas. Current pangenome efforts either lack base-level resolution or are constrained by genomic anchoring, hindering the exploration of genetic signatures of domestication. Here, we assembled chromosome-level genomes for a mallard (Anas platyrhynchos) and an eastern spot-billed duck (Anas zonorhyncha) that retained substantial domestic components. By integrating these with 14 high-quality Anas assemblies, we constructed a base-level resolution pangenome graph encompassing 243.4 Mb of non-reference sequences. Our graph-based pipeline demonstrated superior mapping performance and identified 67,165 population-level SVs, representing a five-fold increase in detection sensitivity compared to linear-based methods. Notably, SVs explained more genetic variance (PC1: 8.75% vs. 3.25%) and provided finer resolution of local ancestry than single nucleotide polymorphisms (SNPs), highlighting the unique advantage of SVs in resolving duck population structure. Leveraging the graph, we recovered 163 ancestral SVs that are prevalent in distant species but persist at low frequencies in domestic and wild populations. The loss of these ancestral sequences in gravity sensing and cilium assembly likely provided the physiological plasticity necessary for domestication. Moreover, highly divergent exonic SVs exhibited a non-random distribution, preferentially accumulating in UTRs over CDS regions. Specifically, highly divergent SVs located in UTRs of TMEM123 and FAM13A may be associated with viral resistance and the regulation of adipogenesis. Similarly, we identified a highly divergent 75 bp insertion located 3.66 kb upstream of the FAM184B gene, which may be associated with changes in muscle growth during duck domestication. In conclusion, the study establishes a comprehensive base-resolution pangenome resource for the Anas. Our findings reveal that SVs are essential for resolving complex evolutionary histories and suggest that SV-mediated regulatory evolution is an important driver of rapid phenotypic change during duck domestication.

RevDate: 2026-07-07

Mohammad SF, Ali F, M Shynara (2026)

Pangenome-guided immunoinformatics design and in silico characterization of a multi-epitope vaccine candidate against Acinetobacter baumannii with nanoparticle assembly potential.

Scientific reports pii:10.1038/s41598-026-59935-4 [Epub ahead of print].

Acinetobacter baumannii is a critical multidrug-resistant pathogen causing severe healthcare infections with high mortality, yet no licensed vaccine exists. This study aims to identify universally conserved surface antigens through pangenome analysis, predict immunogenic epitopes using integrated machine learning, and computationally design and in silico characterize a self-assembling nanoparticle vaccine with dual adjuvants. A computational framework integrating pangenome analysis of 712 complete genomes, epitope prediction, and structural vaccinology was employed to design a multi-epitope nanoparticle vaccine candidate for experimental evaluation. Pangenome analysis identified 3894 core genes with 42 outer membrane proteins, prioritizing OmpA, BamA, and OmpW as antigen targets. Protein language models predicted conformational B-cell epitopes, while NetMHCpan-4.2 predicted T-cell epitopes across 125 HLA alleles. The final construct (AB-VAX-01, 289 amino acids) incorporates 15 epitopes fused with dual adjuvants (RS09 TLR4 and cGAMP STING agonists) and a foldon domain for nanoparticle assembly. Microsecond molecular dynamics simulations with replicates demonstrated stability with TLR4 and STING. Conservation analysis across all 712 genomes showed 96.2-100% epitope identity. Immune simulations predicted Th1-biased responses with 94.2% global population coverage. In silico cloning confirmed favorable codon adaptation parameters. This in silico characterized vaccine construct represents a promising candidate requiring experimental validation against A. baumannii infections.

RevDate: 2026-07-09

Pierre B, Bacilieri R, This D, et al (2026)

Pangenomic analyses in the cultivated grapevine confirm high genomic collinearity and extensive dispensable gene content likely involved in adaptation.

G3 (Bethesda, Md.) pii:8729008 [Epub ahead of print].

Pangenomes have now been developed for several horticultural crops, yet the extent to which genome diversity in sequence and organization contribute to plant adaptation and major agronomic traits remains poorly understood. Here, we assembled the genomes of nine cultivated grapevine varieties and compared the genomes of 15 cultivated grapevine varieties for variation in gene and TE content. We found that genomic collinearity is highly conserved among varieties. We still observed substantial variation across genomes. Notably, we identified across varieties 55,662 orthologous genes, of which 55.3% appears to be dispensable. Dispensable genes are enriched for functions related to adaptation to biotic and abiotic constraints, suggesting that they may play a role in adaptation. Comparing our results with a recently published study, we found substantial differences with ∼12.6 % of the genes we classified as core genes being classified as dispensable genes in this other study. We then constructed a pangenome graph and used it to performed genome-wide association studies for three important traits in grapevine production, which allowed us to include large structural variants as markers in the analyses. We identified 32 loci that we did not detect when we used the PN40024 genome as a reference, 20 of which are newly reported associations. Overall, our results indicates that despite recent advances in characterizing plant pangenomes, current gene classification into core and dispensable gene categories should be taken with caution. They also highlight the value of incorporating structural variants into GWAS, to better characterize the genetic architecture of agronomic traits.

RevDate: 2026-07-09
CmpDate: 2026-07-10

Hu WS, An SH, Kim DW, et al (2026)

Comparative genomics and phenotypes of Listeria monocytogenes isolated from enoki mushrooms in South Korea and China.

Food microbiology, 140:105182.

Listeria monocytogenes is a gram-positive and facultatively anaerobic foodborne pathogen causing listeriosis. The detection of L. monocytogenes in enoki mushrooms sourced from South Korea and China is particularly critical, given their confirmed implication in recent serious listeriosis outbreaks across the global food supply chain. This study investigates the prevalence, genetic diversity, and phenotypical characteristics of L. monocytogenes in enoki mushrooms from South Korea and China. Out of 129 samples, 24 (18.6%) tested positive, with contamination rates of 17.5% in Korean mushrooms and 19.7% in Chinese mushrooms. Whole-genome sequencing, cgMLST, and pan-genome analyses resolved lineage and sequence-type distributions, revealing predominant serogroup 1/2a (83.3%) and lineage II (90.9%). The pan-genomic assessment of L. monocytogenes strains originating from diverse geographical locations indicated the presence of open genomes, which establishes a strong genetic underpinning for adaptation to varied environments. These strains carrying a multitude of virulence genes that significantly contribute to their heightened pathogenic potential. The phylogenetic tree further demonstrated that these highly related Korean and Chinese isolates were intricately intermingled with outbreak-related strains from the USA, Canada, and Europe, confirming a minimal core-genome genetic distance across the global supply chain. This highly homogenous clone, which was ultimately traced back to enoki mushrooms in South Korea and China, suggests that the globalization of the food trade is the primary driver of its rapid, international dissemination.

RevDate: 2026-07-10

Lu T, Li C, Wei H, et al (2026)

NTM-DB: A Comprehensive Non-tuberculosis Mycobacteria Genomic Database.

Genomics, proteomics & bioinformatics pii:8729461 [Epub ahead of print].

Non-tuberculous mycobacteria (NTM) are a major group of environmental bacteria, approximately one-third of which cause serious human infections, particularly respiratory diseases. The global rise in the prevalence and severity of NTM infections has posed a major public health challenge. While high-throughput sequencing has generated vast genomic data on NTM, there remains a lack of comprehensive resources for cross-species genomic analysis. To address these limitations, we developed a specialized database, the Non-tuberculosis Mycobacteria Genomic Database (NTM-DB), tailored for NTM researchers and clinicians. NTM-DB offers the most comprehensive collection of NTM genomic and bioinformatic resources, including 16,469 genome assemblies (13,134 newly assembled genomes), 189 type/standard strain genomes representing 177 species and 12 subspecies, 705 multi-locus sequence typing (MLST) types, 33,240 resistance genes, and 74,315 virulence genes. A user-friendly interactive website was constructed to enable efficient browsing, MLST profiling, searching, online analysis, and downloading of the aforementioned data. Notably, with online analysis tools, users can perform customized genotyping, cross-species phylogeny, pan-genome, and virulence and drug resistance gene annotation analyses using our data and/or their uploaded data. Overall, with its comprehensive data, intuitive interface, and powerful analysis tools, NTM-DB serves as an important resource and reference for NTM researchers and clinicians, thereby improving the diagnosis and treatment of various NTM-related diseases and supporting both scientific discovery and clinical practice. NTM-DB is publicly accessible at https://ngdc.cncb.ac.cn/ntmdb.

RevDate: 2026-07-10

Weis KS, Kaur A, Ghosh P, et al (2026)

Genome evolution in plant pathogenic bacteria.

Genome biology and evolution pii:8729594 [Epub ahead of print].

Bacterial plant pathogens have ravaged crops since the dawn of agriculture and continue to pose a serious threat today. Bacteria and their plant hosts have co-evolved in an evolutionary arms race, with artificial selection due to agriculture tipping the scale in favor of the pathogen. This review gives an overview of plant pathogenic bacterial diversity, showing that pathogenicity has independently evolved numerous times, and that there is not one unifying trait determining plant pathogenicity. Instead, these bacteria represent repeated, independent evolutionary transitions driven by life in complex ecological networks, that include plant hosts, insect vectors, microbial competitors, and highly heterogenous abiotic environments. Their genomes reflect this interplay through a dynamic balance of architecture and flux. These structural features, along with highly variable pangenomes, capture the balance between genome stability and flux imposed by ecological constraints and epidemiological dynamics. Horizontal gene transfer via conjugative plasmids, prophages, integrative and conjugative elements, transposons, and in some lineages, natural competence, remains the major source of adaptive novelty, enabling rapid remodeling of virulence repertoires, metabolic capabilities, and antibiotic or heavy metal resistance genes. These changes create distinct selective landscapes. Agricultural practices such as chemical use, host resistance deployment, or seed trade, can drive recurrent bottlenecks, expansions, and admixture events that leave strong genomic signatures in pathogens. Finally, this review explores the genomic differences enabling the divergence of lifestyles, while also acknowledging knowledge gaps and future directions of research on the evolution of bacterial plant pathogens.

RevDate: 2026-07-10
CmpDate: 2026-07-10

Butler G, Ramakrishnan S, Collins T, et al (2026)

Indirect genomic effects shape cancer risk across species.

bioRxiv : the preprint server for biology pii:2026.06.29.735167.

Tumour prevalence varies dramatically throughout the animal kingdom despite broadly conserved cellular and developmental processes, raising the question of how evolution has shaped susceptibility [1,2] . Here, we link macroevolutionary variation in tumour prevalence to gene-level selection by integrating comparative genomics data from 109 species of birds and mammals using a Bayesian phylogenetic framework to estimate pangenome-wide rates of genetic evolution across >150 million years of evolutionary change. We identify 3,206 genes in which natural selection is associated with shifts in tumour prevalence, with more than 80% of which are linked to reduced prevalence, suggesting pervasive selection for cancer suppression. Using causal phylogenetic inference, we show that genes associated with reduced tumour prevalence act predominantly through indirect effects on body size, revealing growth as a key mediator of cancer risk across species. In contrast, genes associated with increased tumour prevalence exert direct effects independent of body size. Finally, at the species-level, we demonstrate that exceptionally low rates of benign tumours do not necessarily coincide with reduced malignancy, revealing that benign and malignant tumour processes are evolutionarily decoupled. Together, these results reveal how natural selection has fine-tuned the link between genotype, phenotype, and cancer risk across species.

RevDate: 2026-07-10
CmpDate: 2026-07-10

Lu S, Liao WW, DeGorter MK, et al (2026)

Pangenome-based human genome analysis improves trait association and genomic prediction.

bioRxiv : the preprint server for biology pii:2026.07.01.735728.

The Human Pangenome Reference Consortium has generated 462 open-access reference genomes and a variation graph that represents differences among them, providing a substrate for pangenome-based analysis methods that overcome the longstanding limitation of comparing all genomic data to a single linear reference. A key unresolved question is the extent to which these approaches can improve trait mapping. We investigate this using the genetics of gene expression variation as a model. We developed a graph-based method (EdgeDepth) for associating sequence variation with traits using short-read genome sequencing data, and show that it captures complex forms of genetic variation missed by other methods. We evaluated trait mapping performance using 430 samples with deep RNA-seq data, and found that pangenomic methods enable the detection of expression quantitative trait loci involving multiallelic indels and structural variants, leading to increased power at a subset of genes. These include 812 genes (7.9% of total) with ≥20% improvement in statistical significance relative to the 1000 Genomes Project callset, and 185 (1.8%) with a 50% improvement, 10 of which are candidates to explain prior GWAS results. Notably, these analyses implicate GBAP1 pseudogene copy number as a causal factor in Crohn's disease, likely via miRNA-mediated regulation of GBA1 , which explains prior GWAS results based on flanking SNPs. The inclusion of pangenome-specific variation also improved the performance of gene expression prediction models, with median variance explained increasing from 10.1% to 12.5%, and 14.6% of genes showing significant improvement (Δr [2] >0.05). Taken together, these results suggest that integration of pangenomic methods into human genetic studies will improve trait association and genomic prediction at a meaningful subset of genes.

RevDate: 2026-07-07
CmpDate: 2026-07-07

Wang MX, Kille B, Nute MG, et al (2026)

Seqwin: ultrafast identification of signature sequences in microbial genomes.

Bioinformatics (Oxford, England), 42(Supplement_1):.

MOTIVATION: Polymerase chain reaction (PCR) enables rapid, cost-effective diagnostics but requires prior identification of genomic regions that allow sensitive and specific detection of target microbial groups, herein referred to as microbial signature sequences. We introduce Seqwin, an open-source framework designed to automate microbial genome signature discovery. Tens of thousands of microbial genomes are now available for a single species, limiting the application of existing manual and automated approaches for identifying signatures. Modern approaches that are capable of leveraging all available microbial genomes will ensure sensitive and accurate DNA signature identification and enable robust pathogen detection for clinical, environmental, and public health applications.

RESULTS: Seqwin builds weighted pan-genome minimizer graphs and uses a traversal algorithm to identify signature sequences that occur frequently in target genomes but remain rare in non-targets. Unlike earlier tools that depend on strict presence or absence of sequences, Seqwin accommodates natural sequence variation and scales to very large genome collections. When applied to genomes from C. difficile, M. tuberculosis, and S. enterica, Seqwin recovered more high-quality signatures than alternative methods with lower computational burden. Seqwin's analysis of nearly 15 000 S. enterica genomes yielded over 200 candidate signatures in three minutes. Seqwin provides an open-source solution for the long-standing need for scalable microbial signature discovery and diagnostic assay design.

Seqwin is available on GitHub (https://github.com/treangenlab/Seqwin) and can be installed via Bioconda (https://bioconda.github.io/recipes/seqwin/README.html). Benchmarking datasets, outputs, and scripts are available on Zenodo (https://doi.org/10.5281/zenodo.19874011).

RevDate: 2026-07-07
CmpDate: 2026-07-07

Sanaullah A, Brown NK, Shakya P, et al (2026)

RLBWT-based LCP computation in compressed space for terabase-scale pangenome analysis.

Bioinformatics (Oxford, England), 42(Supplement_1):.

MOTIVATION: Lossless full text indexes are utilized in a myriad of applications in bioinformatics. The continuously decreasing cost of generating biological data has resulted in the need to build full text indexes on biological datasets of increasing size. Many compressed full text indexes have been developed to address this problem. In particular, run-length Burrows-Wheeler transform (RLBWT) based compressed full text indexes have seen wide development and adoption. However, the construction of these RLBWT-based compressed full text indexes is still computationally expensive, sometimes prohibitively so, even for current dataset sizes.

RESULTS: Therefore, we present algorithms for the construction of RLBWT-based compressed full text indexes and their supporting data structures in compressed space. The algorithms have a space complexity of O(r) words and run in O(n) time for repetitive datasets, where r is the number of runs in the BWT, n is the length of the text, and repetitive datasets implies nr∈Ω(log n). We provide the first algorithm to compute LCP-related information for repetitive datasets in optimal time and O(r) space, greatly reducing memory requirements. The key idea behind this algorithm is the utilization of r samples of the inverse suffix array at regular intervals. For example, on the Human Pangenome Reference Consortium Release 2 dataset, this reduces peak memory from 2135 GiB to 170 GiB (12.6x reduction) compared to the previous best method (pfp-thresholds).

The implementation is available at https://github.com/ucfcbb/TeraTools.

RevDate: 2026-07-06
CmpDate: 2026-07-06

Soto-Serrano A, Vincze T, Roberts RJ, et al (2026)

Comparative genomics and methylome profiling of Pseudolactococcus laudensis reveal signatures of niche adaptation and strain-level variation in mobile genetic elements and phage defence.

Microbial genomics, 12(7):.

Pseudolactococcus laudensis (formerly named Lactococcus laudensis) is an emerging lactic acid bacterium first isolated from raw milk in 2015 and subsequently detected in vegetables and dairy mesophilic starter cultures. Despite its recurrent isolation from diverse environments, the genetic basis of its niche adaptation, horizontal gene transfer and phage defence remains unexplored. Here, we perform the first comparative genomic and epigenomic analysis of P. laudensis using complete genomes of a plant-derived isolate (MCRI-603), a milk isolate (DSM 28961) and 20 strains from a Danish dairy mesophilic starter culture. Genomes were annotated and analysed using pangenomics, Clustering of Orthologous Genes and methylome profiling. Average nucleotide identity, pangenome and Clustering of Orthologous Genes analyses revealed niche-associated structure: dairy starter strains formed a tight cluster, while the plant isolate MCRI-603 and milk isolate DSM 28961 were more similar to each other than to the starter culture group. The pangenome comprised 4,946 genes, with 1,396 core genes. Dairy starter strains showed markedly elevated numbers of insertion sequences, pseudogenes, plasmids and genomic islands relative to MCRI-603, which was plasmid-free and carried very few insertion sequence elements or genomic islands. DSM 28961 displayed pseudogene count similar to the dairy starter strains but markedly fewer transposases. These patterns are consistent with a plant-associated origin of P. laudensis and progressive dairy specialization via mobile genetic element acquisition. The P. laudensis mobilome was found to carry key niche-related traits. Lactose utilization operons were plasmid-encoded, whereas exopolysaccharide-encoding loci, opp oligopeptide transport systems and several defence loci, including clustered regularly interspaced short palindromic repeats and CRISPR-associated proteins (CRISPR-Cas), were consistently encoded within chromosomal integrative elements. All strains harboured prophage-like elements, including putatively intact prophages in 13 of them, and ~67% of 238 predicted antiphage systems resided on mobile genetic elements, underscoring their central role in phage defence. Restriction-modification systems dominated the defensome, and three strains encoded CRISPR-Cas systems (including type III-A and type I-C), indicating a higher prevalence than has been reported for Lactococcus lactis and Lactococcus cremoris, where CRISPR-Cas has rarely been observed. Methylome analysis identified 43 distinct motifs, of which 25 were novel. The P. laudensis methylome was overwhelmingly dominated by N[6]-methyladenine, and most motifs were short, non-palindromic and largely associated with type III restriction-modification systems and some type I and II subtypes. Nearly all strains exhibited distinct methylation profiles, including those isolated from the same dairy starter culture, highlighting extensive epigenetic diversification in dairy environments. Altogether, the data reveals a highly dynamic genomic and epigenomic landscape in P. laudensis, greatly shaped by mobile genetic elements, and provides a foundation for future work in this species and other Pseudolactococci.

LOAD NEXT 100 CITATIONS

ESP Quick Facts

ESP Origins

In the early 1990's, Robert Robbins was a faculty member at Johns Hopkins, where he directed the informatics core of GDB — the human gene-mapping database of the international human genome project. To share papers with colleagues around the world, he set up a small paper-sharing section on his personal web page. This small project evolved into The Electronic Scholarly Publishing Project.

ESP Support

In 1995, Robbins became the VP/IT of the Fred Hutchinson Cancer Research Center in Seattle, WA. Soon after arriving in Seattle, Robbins secured funding, through the ELSI component of the US Human Genome Project, to create the original ESP.ORG web site, with the formal goal of providing free, world-wide access to the literature of classical genetics.

ESP Rationale

Although the methods of molecular biology can seem almost magical to the uninitiated, the original techniques of classical genetics are readily appreciated by one and all: cross individuals that differ in some inherited trait, collect all of the progeny, score their attributes, and propose mechanisms to explain the patterns of inheritance observed.

ESP Goal

In reading the early works of classical genetics, one is drawn, almost inexorably, into ever more complex models, until molecular explanations begin to seem both necessary and natural. At that point, the tools for understanding genome research are at hand. Assisting readers reach this point was the original goal of The Electronic Scholarly Publishing Project.

ESP Usage

Usage of the site grew rapidly and has remained high. Faculty began to use the site for their assigned readings. Other on-line publishers, ranging from The New York Times to Nature referenced ESP materials in their own publications. Nobel laureates (e.g., Joshua Lederberg) regularly used the site and even wrote to suggest changes and improvements.

ESP Content

When the site began, no journals were making their early content available in digital format. As a result, ESP was obliged to digitize classic literature before it could be made available. For many important papers — such as Mendel's original paper or the first genetic map — ESP had to produce entirely new typeset versions of the works, if they were to be available in a high-quality format.

ESP Help

Early support from the DOE component of the Human Genome Project was critically important for getting the ESP project on a firm foundation. Since that funding ended (nearly 20 years ago), the project has been operated as a purely volunteer effort. Anyone wishing to assist in these efforts should send an email to Robbins.

ESP Plans

With the development of methods for adding typeset side notes to PDF files, the ESP project now plans to add annotated versions of some classical papers to its holdings. We also plan to add new reference and pedagogical material. We have already started providing regularly updated, comprehensive bibliographies to the ESP.ORG site.

Electronic Scholarly Publishing
961 Red Tail Lane
Bellingham, WA 98226

E-mail: RJR8222 @ gmail.com

Papers in Classical Genetics

The ESP began as an effort to share a handful of key papers from the early days of classical genetics. Now the collection has grown to include hundreds of papers, in full-text format.

Digital Books

Along with papers on classical genetics, ESP offers a collection of full-text digital books, including many works by Darwin and even a collection of poetry — Chicago Poems by Carl Sandburg.

Timelines

ESP now offers a large collection of user-selected side-by-side timelines (e.g., all science vs. all other categories, or arts and culture vs. world history), designed to provide a comparative context for appreciating world events.

Biographies

Biographical information about many key scientists (e.g., Walter Sutton).

Selected Bibliographies

Bibliographies on several topics of potential interest to the ESP community are automatically maintained and generated on the ESP site.

ESP Picks from Around the Web (updated 28 JUL 2024 )