Gene Filtering Strategies For Machine Learning-Guided Biomarker Discovery In Neonatal Sepsis RNA-Seq Data.pdf

fgene-14-1158352.pdf
Preview of Gene Filtering Strategies for Machine Learning-Guided Biomarker Discovery in Neonatal Sepsis RNA-Seq Data
🔗 Source: orca.cardiff.ac.uk
📊 Size: 1.07 MB
📄 Pages: 12 pages
⬇️ Downloads: 364

Summary

Supervised machine learning (ML) analysis of gene expression data is a widely used AI approach to derive informative gene subsets with potential utility as diagnostic and prognostic clinical biomarkers. These ML-based feature selection algorithms have the power to both eliminate redundant genes, and identify relevant genes that both discriminate between clinical conditions of interest, and that generalise over the relevant patient populations. To date, the majority of validated sepsis biomarkers have been derived from microarray gene expression data. In recent years RNA-Seq has largely replaced microarray as the preferred technology for generating gene expression data, given the higher resolution and decreasing cost of sequencing. The combination of ML gene selection with richer RNA-Seq data presents the opportunity to discover previously unidentified sepsis biomarkers. ML classification algorithms are however highly sensitive to any feature data characteristic, regardless of scale, that may differ between experimental groups, and will exploit these data characteristic differences in gene selection. Systematically adding noise to feature data has been shown to significantly degrade classification performance across a range of ML classification algorithms. It is therefore desirable to minimise this interference to ensure patterns detected have biological relevancy. To date the impact of applying independent filtering and pre-processing strategies to eliminate noise on the performance of ML approaches to biomarker discovery with RNA-Seq data has not been fully investigated. Here we consider two sources of noise inherent in RNA-Seq data that may negatively impact ML gene selection. Firstly low read counts. Genes with consistently low read count values across all replicates may be technical or biological stochastic artefacts such as the detection of a transcript from a gene that is not uniformly active in a heterogeneous cell population or as the result of a transcriptional error. Below some count threshold, genes with low read counts are subject to greater dispersion (variability) with greater false negatives (zero inflation) and false positives (outliers) that are not representative of true biological differences related to the condition of interest. The filtering out of low read count genes from RNA-Seq data in differential expression analyses is reported to improve detection of differentially expressed genes by reducing the impact of multiple testing corrections. A wide variety of approaches to filtering low count genes have been proposed, the most common being maximum-based filters, where genes with a maximum normalised count over all samples below a threshold t are filtered out. At a lower threshold t, the number of genes retained after filtering is expected to be more variable between samples, given differences in low count noise between samples. As t increases, the number of genes retained is expected to converge as the read counts increasingly represent the true biological signal, which is more consistent across samples. To avoid setting arbitrary t, we propose a systematic objective strategy for removal of uninformative and potentially biasing biomarkers representing up to 60% of transcripts in different sample size datasets.

Description

Researchers developed gene filtering strategies to identify biomarkers for neonatal sepsis using machine learning and RNA-Seq data.

Technical Information

  • File Format: PDF
  • File Size: 1.07 MB
  • Pages: 12
  • Language: EN
  • Total Downloads: 364
  • Last Updated: 1 week ago

Document Overview

This PDF document about Gene Filtering Strategies for Machine Learning-Guided Biomarker Discovery in Neonatal Sepsis RNA-Seq Data provides comprehensive information and guidance. Whether you're a beginner or advanced user, this resource offers valuable insights into Gene Filtering Strategies for Machine Learning-Guided Biomarker Discovery in Neonatal Sepsis RNA-Seq Data.

Related Topics

If you're interested in Gene Filtering Strategies for Machine Learning-Guided Biomarker Discovery in Neonatal Sepsis RNA-Seq Data, you might also want to explore:

Download Gene Filtering Strategies for Machine Learning-Guided Biomarker Discovery in Neonatal Sepsis RNA-Seq Data eBooks for free and learn more about Gene Filtering Strategies for Machine Learning-Guided Biomarker Discovery in Neonatal Sepsis RNA-Seq Data. These books contain exercises and tutorials to improve your practical skills, at all levels!

Not satisfied with this document? We have related documents to Gene Filtering Strategies for Machine Learning-Guided Biomarker Discovery in Neonatal Sepsis RNA-Seq Data, try searching with similar keywords: Gene Filtering Strategies for Machine Learning-Guided Biomarker Discovery in Neonatal Sepsis RNA-Seq Data, Mda Maine Seq Rrbs Seq Ribo Seq Illumina, Neonatal Sepsis Caused By Gram Negative Bacteria In A Neonatal, Biomarker Sepsis, "Hinder och möjliggörare för 1.5°-livsstilar: Ytliga och djupgående strukturella faktorer som påverkar potentialen för hållbar k, Ändring av genomföranderam för en europeisk plattform för utbyte av balansenergi från frekvensåterställn ingsreserver med manuell, Rekommendationer för vaccination mot covid-19 för särskilda grupper av barn -, förstudie för att utvärdera förutsättningarna att genom en innovationsupphandli ng utveckla en drifttjänst för geoenergilager

You can download PDF versions of the user's guide, manuals and ebooks about Gene Filtering Strategies for Machine Learning-Guided Biomarker Discovery in Neonatal Sepsis RNA-Seq Data, you can also find and download for free A free online manual (notices) with beginner and intermediate, Downloads Documentation, You can download PDF files (or DOC and PPT) about Gene Filtering Strategies for Machine Learning-Guided Biomarker Discovery in Neonatal Sepsis RNA-Seq Data for free, but please respect copyrighted ebooks.