Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (229)

Search Parameters:
Keywords = k-mer

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
34 pages, 6903 KB  
Article
Mutation-Aware Machine Learning Framework for Predicting Binding Affinity of Nirmatrelvir Analogs Targeting Coronavirus Main Proteases
by Md Saidur Rahman, Md Mehedi Hasan and Shahidul M. Islam
Molecules 2026, 31(17), 2949; https://doi.org/10.3390/molecules31172949 (registering DOI) - 22 Aug 2026
Abstract
The emergence of resistance-associated mutations in coronavirus main protease (Mpro) poses a significant challenge to the development of broad-spectrum antiviral therapeutics. In this study, we improved and accelerated a mutation-aware machine learning (ML) framework to predict the binding score of Nirmatrelvir analogue ligands [...] Read more.
The emergence of resistance-associated mutations in coronavirus main protease (Mpro) poses a significant challenge to the development of broad-spectrum antiviral therapeutics. In this study, we improved and accelerated a mutation-aware machine learning (ML) framework to predict the binding score of Nirmatrelvir analogue ligands against wild-type and mutant MERS-CoV Mpro. A library of 15,889 Nirmatrelvir derivatives generated through systematic scaffold modification was docked against the wild-type and five variants of the Mpro, producing a total of 95,334 structural and docking score datasets of these protein–ligand complexes. During the ML model development phase, ligand effects were learned from RDKit molecular descriptors and graph-based representations, and the mutation-induced effects were captured through delta-encoded physicochemical properties (hydrophobicity, charge, aromaticity, and polarity) of the active-site residues. Among the evaluated models, the CatBoost regressor tree-based algorithm achieved the lowest mean absolute error (MAE) value of 0.23 Kcal/mol and an R2 of 0.87. Further improvement was achieved by creating a weighted ensemble model combining the CatBoost regressor, XGBoost and LightGBM regressor, resulting in a prediction accuracy with a MAE of 0.19 Kcal/mol and an R2 of 0.90 relative to docking scores. Model robustness was further evaluated through random-, ligand group- and scaffold group- K-fold cross-validation along with their Y-randomization. Moreover, the models were also tested with a new set of 1000 structurally diverse compounds. SHAP analysis was conducted, which identified 20 molecular descriptors critical for accurate predictions. The ensemble model accurately predicted the binding affinities of Nirmatrelvir and its four analogues (E1–E4), reproducing the experimental pIC50 trend and correctly identifying the most potent inhibitors. The ensemble model also showed consistent performance across all MERS-CoV Mpro variants, S147Y, S142G, L144A, S142G/S147Y, and S142G/L144A/S147Y, demonstrating its potential for rapidly discovering mutation-resistant antiviral drugs. Full article
(This article belongs to the Special Issue Computational Approaches for Drug and Protein Design)
15 pages, 12480 KB  
Article
A Cryptanalytic Test of the Phaistos Disc as a Protein Sequence
by Federico M. Giorgi and Fabrizio Baldacci
Cryptography 2026, 10(4), 60; https://doi.org/10.3390/cryptography10040060 - 19 Aug 2026
Viewed by 151
Abstract
The Phaistos Disc is an undeciphered Minoan-era artifact bearing 242 stamped occurrences of 45 distinct signs. We test whether it encodes one or more protein sequences, with each sign representing an amino acid or post-translationally modified amino acid and the 45 signs mapping [...] Read more.
The Phaistos Disc is an undeciphered Minoan-era artifact bearing 242 stamped occurrences of 45 distinct signs. We test whether it encodes one or more protein sequences, with each sign representing an amino acid or post-translationally modified amino acid and the 45 signs mapping many-to-one onto the 20 canonical residues. We treat this as a substitution cipher using a pre-specified battery of methods. Map-invariant statistics disfavor the hypothesis: the disc is more repetitive than length-matched proteins, with a 15-sign exact repeat on Side A, and one sign occurs 19 times, always in group-initial position, a confinement not observed for any amino acid. A Shannon unicity-distance calculation on Swiss-Prot shows that 242 signs are too few to identify a unique map from k-mer statistics. Simulated annealing fails to make the disc more protein-like than optimizer-matched nulls, and a known-plaintext attack against Swiss-Prot finds no decoding that beats a shuffled-disc null in any reading direction. Short stretches can be read as fragments of real proteins, from a woolly mammoth enzyme to a human topoisomerase, but the null predicts such matches by chance. We claim no decipherment and report effect sizes and p-values against matched nulls; the protein hypothesis is not supported. Full article
Show Figures

Figure 1

20 pages, 1685 KB  
Article
Identification of Partitivirus-like RdRPs in the Brevipalpus yothersi Genome Supports Viral-to-Arthropod Horizontal Gene Transfer
by Bruno Afonso Corrêa, Thaís Medinilha Pancher, Denis Calandriello Calio, Aline Daniele Tassi, Laura Rossetto Pereira, Daniel Carrillo, Ricardo Harakava, Valdenice Moreira Novelli, Elliot Watanabe Kitajima, Pedro Luis Ramos-González, Juliana Freitas-Astúa and Daniel Gonzalez-Ibeas
Viruses 2026, 18(8), 892; https://doi.org/10.3390/v18080892 - 13 Aug 2026
Viewed by 561
Abstract
RNA-dependent RNA polymerases (RdRPs) are essential enzymes involved in RNA virus replication and eukaryotic RNA silencing. They are generally absent in vertebrates but present in some invertebrate lineages, such as nematodes and certain arthropods. Brevipalpus yothersi is a phytophagous mite of agricultural relevance [...] Read more.
RNA-dependent RNA polymerases (RdRPs) are essential enzymes involved in RNA virus replication and eukaryotic RNA silencing. They are generally absent in vertebrates but present in some invertebrate lineages, such as nematodes and certain arthropods. Brevipalpus yothersi is a phytophagous mite of agricultural relevance due to its role as a vector of plant-infecting viruses. We have identified two RdRPs on the genome of this mite species that, unexpectedly, are not of eukaryotic origin. Phylogenetic reconstruction and comparisons of 3D protein structures revealed similarity with viral RdRPs of the Partitiviridae family. Both RdRPs retain conserved catalytic motifs at the protein sequence level, and expression was confirmed by RNAseq and qPCR across mite developmental stages, with a peak during the nymphal stage. K-mer profiles showed similarity with mite endogenous genes, suggesting gene amelioration after the integration, or being derived from a viral donor already adapted to the mite. Our study also identified orthologs in other Brevipalpus species, but not in other Acari relatives, supporting that the horizontal gene transfer event is circumscribed to the Brevipalpus genus. These findings highlight an intriguing case of viral gene domestication in arthropods that might influence their developmental biology and the host–virus interaction. Full article
(This article belongs to the Section Invertebrate Viruses)
Show Figures

Graphical abstract

12 pages, 1840 KB  
Article
An Interpretable Multi-Objective Machine Learning Framework for In Silico Prioritization of Anti-Staphylococcus aureus Antimicrobial Peptides
by Jianguo Xu, Donghua Yang, Qingyong Zheng, Tengfei Li, Yating Cui and Jinhui Tian
Diagnostics 2026, 16(15), 2430; https://doi.org/10.3390/diagnostics16152430 - 1 Aug 2026
Viewed by 309
Abstract
Background: Staphylococcus aureus, including methicillin-resistant lineages, is a leading cause of device- and catheter-related infection, and rising resistance motivates the search for antimicrobial peptides (AMPs) with strong anti-staphylococcal activity and low host toxicity. Machine learning can prioritize candidate peptides. However, the [...] Read more.
Background: Staphylococcus aureus, including methicillin-resistant lineages, is a leading cause of device- and catheter-related infection, and rising resistance motivates the search for antimicrobial peptides (AMPs) with strong anti-staphylococcal activity and low host toxicity. Machine learning can prioritize candidate peptides. However, the literature-derived AMP datasets are prone to homology-driven optimism, and computational studies frequently overstate their translational reach. Methods: We curated 4007 deduplicated S. aureus-active AMP records and 582 binary-labeled hemolysis records. Each peptide was encoded with a transparent 538-dimensional physicochemical and compositional feature vector. Five classifiers and five regressors were evaluated for four endpoints (potency classification, log10 MIC regression, hemolysis classification, normalized hemolytic index) under both conventional random 5-fold cross-validation and homology-aware cross-validation, in which sequences were clustered by 3-mer similarity and whole clusters were confined to single folds. Class imbalance was handled by class weighting. Model behavior was interpreted with SHAP and Fisher-exact k-mer enrichment, and candidates were ranked by a multi-objective score that combines the independently trained heads. Results: Under homology-aware validation, performance was lower than under random splitting, as expected. Potency classification reached an AUROC of about 0.71 (Random Forest), compared with 0.797 under random cross-validation. Hemolysis classification remained strong at AUROC 0.90 (95% CI 0.88 to 0.93), which indicates that its high accuracy is not a homology leakage artifact. MIC regression was modest (homology-aware R2 0.17, Spearman ρ 0.39) and is therefore treated only as a rank-ordering signal. SHAP and k-mer analyses recovered interpretable structure–activity relationships. Net positive charge and amphipathicity drove potency, whereas bulk hydrophobicity drove hemolysis. Applying the pipeline to a generated pool prioritized 20 candidates. Nearest-neighbor analysis shows that these are close optimized variants of known potent scaffolds, with a median identity of 95% to a known peptide, rather than novel sequences. Conclusions: We present an interpretable, honestly benchmarked multi-objective pipeline that optimizes known anti-S. aureus AMP scaffolds toward lower predicted hemolysis. The prioritized peptides are computational hypotheses for future synthesis and experimental testing. Their low predicted hemolysis reflects a selection criterion rather than validated safety, and cross-species selectivity was not assessed. Full article
Show Figures

Figure 1

18 pages, 3186 KB  
Article
A Haplotype-Resolved Genome Assembly of the Long-Spined Sea Urchin Diadema antillarum, a Keystone Caribbean Reef Herbivore
by Audrey J. Majeske, Juliet M. Wong, Carlos A. Farkas Pool, Jose M. Eirin-Lopez, Jose V. Lopez, Walter Wolfsberger, Nikolaos V. Schizas, Alondra M. Díaz-Lameiro, Stephanie O. Castro-Márquez, Kenneth Hilkert, Alejandro J. Mercado Capote and Taras K. Oleksyk
Genes 2026, 17(8), 876; https://doi.org/10.3390/genes17080876 - 28 Jul 2026
Viewed by 589
Abstract
Background/Objectives: The long-spined sea urchin Diadema antillarum is a keystone herbivore whose grazing maintains Caribbean coral reefs; basin-wide mass mortalities in 1983–1984 and 2022 have made genomic resources a conservation priority, yet no nuclear genome existed for the species. We aimed to [...] Read more.
Background/Objectives: The long-spined sea urchin Diadema antillarum is a keystone herbivore whose grazing maintains Caribbean coral reefs; basin-wide mass mortalities in 1983–1984 and 2022 have made genomic resources a conservation priority, yet no nuclear genome existed for the species. We aimed to generate the first nuclear reference and to resolve the high heterozygosity that complicates genome assembly in broadcast-spawning marine invertebrates. Methods: For the assembly, we combined PacBio HiFi, Oxford Nanopore, and Illumina sequencing. Genome size and heterozygosity were estimated by k-mer profiling. We compared standard and haplotype-aware assembly strategies (hifiasm), evaluated completeness with BUSCO, and annotated repeats using a species-specific RepeatModeler library. Results: k-mer profiling estimated a haploid genome of ~703 Mb with 2.52% heterozygosity. Standard assembly then produced an inflated 1.75 Gb assembly (98.4% BUSCO-complete but 84.4% duplicated), indicating retention of both haplotypes. Haplotype-aware reassembly separated this into a collapsed primary assembly (1.03 Gb) and two phased haplotypes (0.95 and 0.89 Gb), each comparable in size to the chromosome-level congener D. antillarum (886 Mb). BUSCO completeness reached 99.0%, with single-copy orthologs rising to 85–90%, and reference-free consensus quality reached QV 44.5 (Merqury; initial assembly). This genome is repeat-rich (42.84% repetitive; 29.96% unclassified). Conclusions: We provide the collapsed primary assembly together with both phased haplotypes as a haplotype-resolved reference for D. antillarum, establishing a foundation for immunogenomic, comparative, and population-genetic studies and for monitoring and restoration of this ecologically critical species. More broadly, the study shows that haplotype-aware assembly is essential for resolving such highly heterozygous genomes and delivers the genomic foundation needed to guide the conservation of this keystone Caribbean reef species. Full article
(This article belongs to the Section Animal Genetics and Genomics)
Show Figures

Figure 1

14 pages, 1962 KB  
Article
First-Trimester Maternal Serum Growth Arrest-Specific Protein 6 (Gas6) Levels for the Prediction of Preeclampsia
by Zehra Yılmaz, Nazan Yurtcu, Dilay Karademir, Canan Soyer Calıskan and Samettin Celik
J. Clin. Med. 2026, 15(14), 5604; https://doi.org/10.3390/jcm15145604 - 17 Jul 2026
Viewed by 390
Abstract
Objective: Preeclampsia (PE) is a leading cause of maternal and perinatal morbidity; however, reliable first-trimester biomarkers are limited. Growth arrest-specific protein 6 (Gas6), a vitamin K-dependent ligand for the TAM receptor family (Tyro3, AXL, and MerTK), is crucial for trophoblast invasion and [...] Read more.
Objective: Preeclampsia (PE) is a leading cause of maternal and perinatal morbidity; however, reliable first-trimester biomarkers are limited. Growth arrest-specific protein 6 (Gas6), a vitamin K-dependent ligand for the TAM receptor family (Tyro3, AXL, and MerTK), is crucial for trophoblast invasion and placental remodeling. This study evaluated whether first-trimester serum Gas6 levels differed between women with PE and controls and assessed the predictive utility across phenotypes. Methods: Participants were enrolled prospectively during routine first-trimester screening at 11 + 0 to 13 + 6 weeks, and cases and controls were selected using a nested case–control design after pregnancy outcomes were known. ROC-derived cut-off values for Gas6 were selected retrospectively using the Youden index and were considered exploratory. The PE group was divided into early (n = 15, <34 weeks) and late onset (n = 65, ≥34 weeks) subgroups. ELISA measured serum Gas6. Mann–Whitney U, Kruskal–Wallis, Spearman’s correlation, ROC analysis, and multivariate logistic regression were applied. Results: Gas6 levels were significantly lower in PE than controls (15.73 vs. 38.08 ng/mL, p < 0.001); control reference median (38.08 ng/mL, IQR 27.62 to 49.14), decreasing progressively to late (20.07 ng/mL) and early onset PE (6.45 ng/mL; all p < 0.001). ROC analysis showed the highest exploratory discrimination for early onset preeclampsia (AUC = 0.959, 95% CI: 0.923–0.995; sensitivity 93.3%, specificity 84.1% at a retrospectively selected cut-off of ≤13.82 ng/mL), with more modest discrimination for overall PE (AUC = 0.762) and late onset PE (AUC = 0.708). Lower first-trimester Gas6 levels remained significantly associated with early onset PE after adjustment for maternal age and BMI (OR = 0.555, 95% CI: 0.412–0.748, p < 0.001). Conclusions: Serum Gas6 levels in the first trimester are lower in PE, particularly in cases of early onset disease, highlighting its potential as a candidate biomarker that requires further prospective and external validation before any screening application can be considered. Full article
(This article belongs to the Section Obstetrics & Gynecology)
Show Figures

Figure 1

49 pages, 2170 KB  
Article
A DNA-Local, Constraint-Aware Dual-Head Transformer for Pseudorandom Stream Generation
by Alev Kaya and İbrahim Türkoğlu
Entropy 2026, 28(6), 694; https://doi.org/10.3390/e28060694 - 16 Jun 2026
Viewed by 395
Abstract
Pseudorandom number generators (PRNGs) used in deoxyribonucleic acid (DNA)-oriented computational workflows often generate outputs in the bit domain and then map them to DNA symbols. This indirect strategy may treat DNA-specific constraints, including GC balance, homopolymer limits, and short-range sequence dependencies, as separate [...] Read more.
Pseudorandom number generators (PRNGs) used in deoxyribonucleic acid (DNA)-oriented computational workflows often generate outputs in the bit domain and then map them to DNA symbols. This indirect strategy may treat DNA-specific constraints, including GC balance, homopolymer limits, and short-range sequence dependencies, as separate from generation. This study proposes a constraint-aware, dual-head decoder-only Transformer framework for DNA-local PRNG generation directly in the adenine/cytosine/guanine/thymine (A/C/G/T) alphabet. The model generates the next DNA base and derives the bitstream through dynamic selection among eight equivalent DNA-to-bit coding rules. The framework was evaluated under R1 based on real genomic data, R1-ext as independent validation, R2 based on synthetic data, and R3 without training or reference data. For each setting, 10 independent runs were performed, each producing a 500,000-base DNA sequence and a 1,000,000-bit stream. Bit-level evaluation used NIST SP 800-22, SP 800-90B-inspired min-entropy/health indicators, and ENT, while DNA-level evaluation used GC balance, homopolymer control, and symbolic structural metrics. The reported NIST tests satisfied the acceptance criterion, t-tuple min-entropy lower bounds ranged from 0.9955 to 0.9964 bit/bit, and core DNA-compatibility constraints were preserved. Multi-stream and exact-match k-mer leakage analyses indicated no systematic bit-level dependence or direct long-fragment copying. Overall, the framework supports reproducible DNA-local PRNG generation and multilayer validation. Full article
(This article belongs to the Section Information Theory, Probability and Statistics)
Show Figures

Figure 1

17 pages, 2565 KB  
Article
Frequency-Domain Transformation of cfDNA End-Motif Profiles Enhances Robust Cancer Detection
by Xinwei Sheng, Xinming Du, Qianqian Shi and Xionghui Zhou
Genes 2026, 17(6), 661; https://doi.org/10.3390/genes17060661 - 5 Jun 2026
Viewed by 747
Abstract
Background/Objectives: Cell-free DNA (cfDNA) end-motifs (EDMs) are promising fragmentomic features for noninvasive cancer detection; however, their diagnostic utility may be limited by background signals from abundant hematopoietic-derived cfDNA fragments. Existing EDM-based approaches, including the Motif Diversity Score (MDS) and classifiers based on [...] Read more.
Background/Objectives: Cell-free DNA (cfDNA) end-motifs (EDMs) are promising fragmentomic features for noninvasive cancer detection; however, their diagnostic utility may be limited by background signals from abundant hematopoietic-derived cfDNA fragments. Existing EDM-based approaches, including the Motif Diversity Score (MDS) and classifiers based on raw motif frequencies, often show limited robustness across different datasets. Methods: To address this limitation, we developed a frequency-domain analytical framework based on the Discrete Fourier Transform (DFT), converting k-mer EDM frequency profiles into amplitude spectral features. We further constructed a stacking-based Ensemble Spectral Model (ESM) integrating multi-scale spectral features from 4–6-mer EDMs. Results: The framework was evaluated using 1782 plasma cfDNA samples from four independent studies comprising six datasets. Raw EDM profiles showed extremely high similarity between cancer and non-cancer samples (mean Spearman R = 0.999). Following DFT transformation, amplitude spectra showed improved separability between groups. Across datasets, the ESM achieved a mean AUC of 0.843, representing a 15.0% improvement over raw 4-mer EDM-based SVM models and a 56.4% improvement over the MDS. At 95% specificity, mean sensitivity reached 0.585, exceeding those of the raw EDM (0.418) and MDS (0.195). Frequency-guided motif attribution further linked spectral features to sequence-level motif patterns and potential regulatory programs. Conclusions: Frequency-domain transformation improves the representation of cfDNA EDM profiles and provides a robust analytical framework for cross-dataset cancer detection. Full article
(This article belongs to the Section Bioinformatics)
Show Figures

Figure 1

16 pages, 2409 KB  
Article
Unsupervised Reference Modeling of Nanopore Signals for DNA/RNA Modification Detection
by Yongji Zou, Mian Umair Ahsan and Kai Wang
Genes 2026, 17(5), 525; https://doi.org/10.3390/genes17050525 - 29 Apr 2026
Viewed by 818
Abstract
Background: Nanopore sequencing produces ionic current signals that are sensitive to chemical modifications in DNA and RNA molecules. However, accurate modification detection remains challenging due to limited labeled data and variability across experimental conditions. Methods: We present a scalable unsupervised framework for modification [...] Read more.
Background: Nanopore sequencing produces ionic current signals that are sensitive to chemical modifications in DNA and RNA molecules. However, accurate modification detection remains challenging due to limited labeled data and variability across experimental conditions. Methods: We present a scalable unsupervised framework for modification discovery that learns reference signal distributions from unmodified sequences using a CNN–Transformer variational autoencoder (VAE). The model is trained on large-scale data via streaming sampling and k-mer-aware soft balancing to ensure robust signal representation. At inference, candidate nucleotides are scored using the VAE reconstruction error, and read-level signals are aggregated to produce site-level modification evidence. Results: On controlled DNA oligonucleotide datasets, models trained on unmodified sequences achieve strong discrimination when evaluated on modified oligos. In contrast, performance decreases in cell line samples when models trained on unmodified whole-genome-amplified (WGA) DNA and in vitro-transcribed (IVT) RNA are evaluated on natively modified (5mC/m6A) data, reflecting the impacts of biological noise and heterogeneity. Despite reduced classification accuracy, site-level anomaly score profiles exhibit peak-like patterns that correspond to known modification-enriched regions. Conclusions: These findings demonstrate the feasibility of large-scale unsupervised reference modeling for de novo modification detection, while underscoring the challenges in translating models built from synthetic oligo datasets into robust genome-wide modification detection. Full article
(This article belongs to the Section Bioinformatics)
Show Figures

Figure 1

17 pages, 7393 KB  
Article
Deciphering 6-mer Spectra Distribution Rules in Coronavirus Genomes: Application to Comparative Genomic Analysis
by Zhenhua Yang, Hong Li, Xiaolong Li and Guojun Liu
Int. J. Mol. Sci. 2026, 27(8), 3604; https://doi.org/10.3390/ijms27083604 - 18 Apr 2026
Viewed by 545
Abstract
Given the rapid mutation and high transmissibility of coronaviruses, especially SARS-CoV-2, comparative genomic studies are crucial for understanding viral evolution, transmission dynamics, and therapeutic development. In prior work, we analyzed and compared the spectral distribution patterns of various k-mer subsets across 920 genome [...] Read more.
Given the rapid mutation and high transmissibility of coronaviruses, especially SARS-CoV-2, comparative genomic studies are crucial for understanding viral evolution, transmission dynamics, and therapeutic development. In prior work, we analyzed and compared the spectral distribution patterns of various k-mer subsets across 920 genome sequences, spanning from primates to prokaryotes. This revealed an evolutionary mechanism in genome sequences, indicating the presence of both CG and TA-specific selection modes. In the present study, we further investigate the specific selection modes in coronavirus genomic sequences by examining the intrinsic distribution rules of 32 XYi 6-mer subset spectra. Our results show that coronavirus genomes exhibit only the CG-specific selection mode, with no evidence of TA-specific selection. Using the CG-specific selection mode, we identified CG1 6-mers as the fundamental subset underlying coronavirus genome evolution. To validate the CG1 subset, we constructed phylogenetic relationships for a set of coronaviruses and SARS-CoV-2 variant genomes. Comparative analysis confirmed that the resulting phylogenetic relationships align more closely with established knowledge. This study thus provides a theoretical framework for inferring phylogenetic relationships at the whole-genome level. Full article
(This article belongs to the Section Molecular Genetics and Genomics)
Show Figures

Figure 1

14 pages, 5203 KB  
Article
Machine Learning Prediction of Listeria monocytogenes Serogroups and Biofilm Formation from Infrared Spectra: A Comparative Study with Genomic Analysis
by Martine Denis, Stéphanie Bougeard, Virginie Allain, Mélanie Guy, Emmanuelle Houard, Arnaud Felten, Jean Lagarde, Benoit Gassilloud, Evelyne Boscher and Pierre-Emmanuel Douarre
Appl. Microbiol. 2026, 6(4), 54; https://doi.org/10.3390/applmicrobiol6040054 - 16 Apr 2026
Viewed by 943
Abstract
This study evaluated the performance of Fourier-transform infrared (FTIR) spectroscopy for identifying spectral signatures associated with two key traits of Listeria monocytogenes: serogroup classification and biofilm-forming capacity. A total of 100 strains, previously serogrouped by PCR and categorized as high, intermediate, or [...] Read more.
This study evaluated the performance of Fourier-transform infrared (FTIR) spectroscopy for identifying spectral signatures associated with two key traits of Listeria monocytogenes: serogroup classification and biofilm-forming capacity. A total of 100 strains, previously serogrouped by PCR and categorized as high, intermediate, or low biofilm producers, were analyzed. Whole-genome sequencing was performed, and comparative genomics was conducted at core-genome, pangenome, and whole-genome (k-mer) levels to determine which genomic representation best reflected the phenotypes. Strains were typed using Fourier-Transform Infrared (FTIR Biotyper® system from Bruker Daltonics GmbH and Co., Bremen, Germany) with five technical replicates. Spectral data from the polysaccharide region (1300–800 cm−1) were extracted and used to train twelve statistical models within a machine learning pipeline combined with cross-validation to predict four serogroups and three biofilm clusters from 501 spectral variables. Genomic analyses showed strong concordance between population structure and serogroup, whereas biofilm formation displayed only weak genomic association, explaining less than 0.1% of genomic variance (PERMANOVA R2 ≤ 0.001). Penalized discriminant analysis achieved the highest performance for serogroup prediction (overall accuracy 97.2%), while the k-nearest neighbor model performed best for biofilm prediction (74.8%). Two dedicated R Shiny applications were developed to facilitate model use. Overall, FTIR spectroscopy coupled with machine learning can provide a rapid and cost-effective alternative to PCR, genomic analyses, and in vitro assays for phenotypic trait prediction. Full article
Show Figures

Figure 1

24 pages, 6873 KB  
Article
Characterisation of Naturally Occurring MERS-CoV Spike Mutations and Their Impact on Fusion and Neutralisation
by Rachael Dempsey, Hannah Goldswain, Joseph Newman, Nazia Thakur, Tracy MacGill, Todd Myers, Robert Orr, Dalan Bailey, James P. Stewart, Waleed Aljabr and Julian A. Hiscox
Viruses 2026, 18(3), 377; https://doi.org/10.3390/v18030377 - 18 Mar 2026
Viewed by 1313
Abstract
In this study, the phenotypic consequences of naturally occurring single nucleotide polymorphisms (SNPs) in the Middle East respiratory syndrome coronavirus (MERS-CoV) Spike protein were investigated. The impact of Spike mutations on the syncytia formation and neutralisation of contemporary MERS-CoV strains is not currently [...] Read more.
In this study, the phenotypic consequences of naturally occurring single nucleotide polymorphisms (SNPs) in the Middle East respiratory syndrome coronavirus (MERS-CoV) Spike protein were investigated. The impact of Spike mutations on the syncytia formation and neutralisation of contemporary MERS-CoV strains is not currently well understood. Mutations were identified by aligning 584 MERS-CoV Spike sequences from either human clinical isolates collected between 2012 and 2024 or from a clinical isolate that had been passaged in human cells. Fifteen SNPs of interest occurring in the N-terminal domain (NTD), receptor binding domain (RBD) and adjacent to the S1/S2 cleavage site were selected for further characterisation based on their location in the Spike protein, frequency and identification in previous studies. A contemporary clade B, lineage 5 wildtype Spike sequence, obtained from a human MERS-CoV clinical isolate, was used as the backbone in this study. The mutations of interest were introduced to the wildtype backbone to generate Spike variants. Spike variants were characterised via cell–cell fusion assays, and a lentiviral pseudotyping system was used to investigate the impact of these Spike mutations on neutralisation. The I529T, E536K and L745F mutations were shown to increase fusion and syncytia formation. The L411F, T424I, L506F, L745F and T746K mutations were found to increase resistance to neutralisation by pooled patient sera. This study has identified novel naturally occurring Spike mutations that resulted in phenotypic differences in the syncytia formation and neutralisation of contemporary MERS-CoV strains. Continued investigation of the phenotypic consequences of MERS-CoV Spike mutations is essential for assessing the risk to public health, especially given the pandemic potential of this virus. Full article
(This article belongs to the Section Coronaviruses)
Show Figures

Figure 1

14 pages, 1128 KB  
Article
Reconstruction of DNA Sequences Through Eulerian Traversal of De Bruijn Graphs
by Baining Zhu, Siqi Liu and Suwei Liu
Mathematics 2026, 14(5), 832; https://doi.org/10.3390/math14050832 - 28 Feb 2026
Viewed by 841
Abstract
Reconstructing a genome from collections of short DNA fragments is a fundamental problem in modern sequencing. Although genome assembly algorithms are widely used in practice, the mathematical conditions that allow exact reconstruction are not always clear. This study develops a graph-theoretic framework for [...] Read more.
Reconstructing a genome from collections of short DNA fragments is a fundamental problem in modern sequencing. Although genome assembly algorithms are widely used in practice, the mathematical conditions that allow exact reconstruction are not always clear. This study develops a graph-theoretic framework for genome reconstruction using De Bruijn graphs and Eulerian paths in an idealized, error-free setting. Each k-mer is represented as a directed edge connecting its (k1)-length prefix and suffix. The resulting overlap graph is constructed using a balanced search tree and traversed with a stack-based Eulerian algorithm. Numerical experiments over a broad range of genome lengths and fragment lengths reveal a sharp transition in reconstruction accuracy. This transition is explained by a probabilistic model for prefix collisions in the directed graph. The theoretical predictions agree with simulation results and provide conditions on the fragment length required for reliable reconstruction. These results show that the difficulty of genome assembly is governed primarily by the combinatorial structure of the underlying graph rather than by algorithmic heuristics. Full article
Show Figures

Figure 1

14 pages, 900 KB  
Article
Alignment-Free Machine Learning Serotype Classification of the Dengue Virus
by Vladimir Gajdov, Isidora Prosic, Mihaela Kavran, Filip Bosilkov, Tamas Petrovic, Jelena Konstantinov and Gospava Lazic
Viruses 2026, 18(3), 280; https://doi.org/10.3390/v18030280 - 25 Feb 2026
Cited by 1 | Viewed by 1358
Abstract
Dengue virus (DENV) serotyping is essential for epidemiological surveillance, clinical risk assessment, and vaccine evaluation, as the four dengue serotypes differ in pathogenicity, immune interactions, and population dynamics. Existing subtyping methods largely rely on sequence alignment and phylogenetic inference, which can be computationally [...] Read more.
Dengue virus (DENV) serotyping is essential for epidemiological surveillance, clinical risk assessment, and vaccine evaluation, as the four dengue serotypes differ in pathogenicity, immune interactions, and population dynamics. Existing subtyping methods largely rely on sequence alignment and phylogenetic inference, which can be computationally intensive and unreliable for short, fragmented, or error-prone sequences commonly generated in diagnostic and surveillance settings. There is a need for fast, alignment-free serotyping approaches that maintain high accuracy across heterogeneous sequence lengths while remaining scalable, transparent, and suitable for real-world diagnostic inputs. We demonstrate that compact 3-mer composition features are sufficient for highly accurate dengue virus serotyping when coupled with a lineage-aware Random Forest classification framework. Using 64 normalized 3-mer frequency features per sequence with ambiguity masking and enforcing strict cluster-aware validation at both 99% and 95% nucleotide identity thresholds, our approach achieved near-perfect accuracy and macro-F1 scores on held-out internal test sets. To further ensure independence, external validation datasets were filtered to remove exact sequence matches and any sequences sharing ≥99% or ≥95% nucleotide identity with internal data. On these strictly independent external datasets, the model maintained 100% accuracy and macro-F1 performance, confirming robust generalization beyond database redundancy. Robustness analyses showed stable performance under contiguous sequence truncation down to 300 bp and in the presence of ambiguous nucleotides, indicating resilience to realistic diagnostic inputs. These results demonstrate that a lightweight, alignment-free, machine learning approach can rival alignment-dependent methods while maintaining strict lineage-aware evaluation controls. The proposed framework combines high predictive accuracy, probabilistic reliability, computational efficiency, and reproducible validation design, making it well suited for large-scale genomic surveillance, rapid pre-screening, and diagnostic decision-support applications. Full article
Show Figures

Figure 1

25 pages, 20668 KB  
Article
Total Saponins from Rhizoma Panacis Majoris Promote Wound Healing in Diabetic Rats by Regulating Inflammatory Dysregulation
by Xiang Xu, Mei-Xia Wang, Ya-Ning Zhu, Xiang-Duo Zuo, Di Hu and Jing-Ping Li
Int. J. Mol. Sci. 2026, 27(2), 955; https://doi.org/10.3390/ijms27020955 - 18 Jan 2026
Cited by 2 | Viewed by 972
Abstract
In individuals with diabetes, dysregulation of inflammatory processes hinders the progression of wounds into the proliferative phase, resulting in chronic, non-healing wounds. Total saponins from Rhizoma Panacis majoris (SRPM), bioactive compounds naturally extracted from the rhizome of Panax japonicus C.A.Mey. var. [...] Read more.
In individuals with diabetes, dysregulation of inflammatory processes hinders the progression of wounds into the proliferative phase, resulting in chronic, non-healing wounds. Total saponins from Rhizoma Panacis majoris (SRPM), bioactive compounds naturally extracted from the rhizome of Panax japonicus C.A.Mey. var. major (Burk.) C.Y.Wu and K.M.Feng, have demonstrated extensive anti-inflammatory and immunomodulatory properties. This study aims to elucidate the molecular mechanisms underlying the facilitative effects of SRPM on diabetic wound healing, with particular emphasis on its anti-inflammatory actions. A high-fat diet combined with streptozotocin (STZ) administration was used to induce type 2 diabetes in rats. After two weeks of oral treatment with SRPM suspension, a wound model was established. Subsequently, a two-week course of combined local and systemic therapy was administered using both SRPM suspension and SRPM gel. SRPM markedly reduces the levels of pro-inflammatory mediators, including IL-1α, IL-1β, IL-6, MIP-1α, TNF-α, and MCP-1, in both rat tissues and serum. Concurrently, it increases the expression of anti-inflammatory cytokines such as IL-10, TGF-β1, and PDGF-BB, while also enhancing the expression of the tissue remodelling marker bFGF. Additionally, SRPM significantly decreases the accumulation of apoptotic cells within tissues by downregulating the pro-apoptotic gene Caspase-3, upregulating the anti-apoptotic gene Bcl-2, and increasing the expression of the apoptotic cell clearance receptor MerTK. Moreover, SRPM inhibits neutrophil infiltration and the release of neutrophil extracellular traps (NETs) in tissues, promotes macrophage polarisation towards the M2 phenotype, and activates the Wnt/β-catenin signalling pathway at the molecular level. SRPM promotes the healing of wounds in diabetic rats potentially due to its anti-inflammatory properties. Full article
(This article belongs to the Section Bioactives and Nutraceuticals)
Show Figures

Figure 1

Back to TopTop