closedPITTSBURGH, PA

Research Framework for Innovation in Read Mapping

National Institute of General Medical Sciences

Description

Read mapping is the central problem of genomic sequence analysis. The problem is to efficiently and correctly align millions of sequencing reads (fragments) to a set of relatively unchanging reference genome sequences. While conceptually a straightforward problem, the challenge comes from the scale of the problem and the frequent need to solve it. The quest to solve it quickly has given rise to new data structures, new algorithms, and highly efficient implementations that are foundational to academic and industry research efforts and to genomic analysis in general. However, while advances in mapping have led to significant reductions in computational resource requirements, it remains one of the most costly steps of most analysis pipelines. Hence, additional improvements in speed and memory are needed. Further, as new technologies and new use cases arise, new methods for mapping must be developed. But this poses a significant challenge: while there are many interesting algorithmic ideas to pursue to improve read mapping, practically testing these ideas requires significant software engineering effort unrelated to the core new ideas. Hence, innovation in read mapping is slower than it needs to be. We will develop a modular read mapper that will serve as a research framework and testbed for innovations in read mapping. This system will be built on a new modular architecture that allows for easily swapping new techniques for various subcomponents. This will be augmented with a flexible build system that allows researchers to distribute new mappers in a lightweight manner without forking or duplicating popular mappers. We will use these module definitions to implement a complete, modern, high-performance modular mapper to serve as a research and innovation platform for read mapping. This mapper will be implemented using literate programming techniques, and we will extend literate programming tools and validate the paradigm for its applicability for creating reproducible bioinformatics software. Using this framework, we will explore many ideas for improving read mapping. These include new deep-learning- based sketching and seeding schemes, new full-text indices, new BAM storage and compression approaches, network-based distributed read mapping, and hyper-parameter and modular selection optimization. These will lead to better mappers that are applicable in new settings and to new approaches for creating extensible research software creation. The proposed project will catalyze greater innovations in read mapping far beyond a single research group. It will improve the tools, processes, and insight in how to structure large-scale, high performance bioinformatics software, and it will itself result in the implementation and validation of various new algorithmic ideas in mapping. Project Number: 1R01GM157795-01A1 | Fiscal Year: 2026 | NIH Institute/Center: National Institute of General Medical Sciences (NIGMS) | Principal Investigator: Carleton Kingsford | Institution: CARNEGIE-MELLON UNIVERSITY, PITTSBURGH, PA | Award Amount: $315,005 | Activity Code: R01 | Study Section: Special Emphasis Panel[ZRG1 BBBT-M (84)] View on NIH RePORTER: https://reporter.nih.gov/project-details/11280671

Interested in this grant?

Start a free 7-day trial to get match scores, save grants, and build your application with AI.

Start free trial

Grant Details

Funding Range

$315,005 - $315,005

Deadline

Not specified

Geographic Scope

PITTSBURGH, PA

Status
closed

View the application link

Start a free 7-day trial to open the original listing and funder website, save this grant, and track its deadline. Cancel anytime.

Start free trial

Want to see how well this grant matches your organization?

Get Your Match Score

Get personalized grant matches

Start your free trial to save opportunities, get AI-powered match scores, and manage your applications in one place.

Start Free Trial