Collaborative Research: A mathematical framework for scalable species tree estimation via site pattern scoring schemes
U.S. National Science FoundationDescription
High-quality genome data are being produced at a pace that is changing evolutionary biology. These data can clarify how species are related, but they also bring a hard problem: evolution is not uniform across the genome. This project develops new mathematics and algorithms for estimating species histories from whole-genome data while allowing a separate history for each region. The approach builds on a recently proposed, fast, and accurate method that uses simple scores computed from patterns in the data but is not yet well understood theoretically. By explaining why its scores work, finding better scores, and extending the method to additional settings, the project will make genome-scale evolutionary analysis more accurate, scalable, and robust. The resulting tools will be distributed as open software and taught through software schools. The project will also train students at the interface of mathematics, computer science, and biology. Its long-term benefits include stronger tools for biological discovery, including work relevant to biotechnology, invasive species, and disease outbreaks, and new ideas for artificial intelligence and machine learning methods for analyzing large heterogeneous datasets. This project will develop a mathematical framework for quartet-based linear scores for species tree estimation from whole-genome alignments. The starting point is CASTER, a site-based method whose empirical accuracy and scalability come from scoring site patterns over quartets and aggregating those scores without enumerating all quartets. The theory of these scores is currently incomplete. The project will characterize valid linear scoring schemes under hierarchical sequence-evolution and gene-tree-evolution models, including the multispecies coalescent, models with multi-copy genes, and substitution-rate heterogeneity. It will analyze the algebraic structure of the score space, connections to phylogenetic invariants, and extensions to site pairs and multi-site patterns. The project will also study the probabilistic properties of score gaps, including their signal-to-noise behavior, in order to design more accurate scores and to develop site-based estimators of branch length, branch support, local histories, and genomic outliers. The resulting algorithms will be implemented by extending CASTER and tested on simulations and large empirical datasets. The work combines applied probability, statistical theory, graph-based algorithms, and algebra, while connecting to machine learning through scalable statistical inference from large heterogeneous genomic data. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria. NSF Award ID: 2601545 | Program: 01002627DB NSF RESEARCH & RELATED ACTIVIT | Principal Investigator: Sebastien Roch | Institution: University of Wisconsin-Madison, MADISON, WI | Award Amount: $327,547 View on NSF Award Search: https://www.nsf.gov/awardsearch/show-award/?AWD_ID=2601545 View on Research.gov: https://www.research.gov/awardapi-service/v1/awards/2601545.html
Interested in this grant?
Start a free 7-day trial to get match scores, save grants, and build your application with AI.
Grant Details
$327,547 - $327,547
Not specified
MADISON, WI
View the application link
Start a free 7-day trial to open the original listing and funder website, save this grant, and track its deadline. Cancel anytime.
Start free trialWant to see how well this grant matches your organization?
Get Your Match Score