RI:Small:Selective Prediction Techniques for Visual-Language Models
U.S. National Science FoundationDescription
Artificial intelligence systems that can analyze images and communicate using natural language are rapidly becoming part of everyday life. These systems are increasingly used in applications such as healthcare, robotics, autonomous vehicles, education, and online information services. However, current visual-language models can produce “hallucinations,” generating statements that sound plausible but are factually incorrect or unsupported by the visual input. In safety-critical settings, these errors can have serious consequences. This project develops trustworthy artificial intelligence systems that can recognize uncertainty, detect potential hallucinations, and avoid making unreliable predictions. The project studies new methods that allow visual-language systems to determine when they should answer a question and when they should abstain due to low confidence. By improving the reliability and safety of artificial intelligence systems, the project advances the national interest through contributions to trustworthy AI, safer deployment of AI technologies, and the training of students in advanced machine learning and computer vision research. This project develops selective prediction techniques for large visual-language models capable of detecting and avoiding hallucinations across a broad range of tasks. The research investigates three complementary directions. First, the project develops new probability calibration algorithms that improve the reliability of model confidence estimates while requiring substantially less training data than existing approaches. Second, the project develops open-vocabulary hallucination verification modules that can be attached to existing visual-language models without retraining, using retrieval-augmented methods and external reference data to assess factual consistency between images and generated text. Third, the project develops scalable training methods for hallucination detection, including automated procedures for large-scale annotation of fine-grained hallucinations and new learning algorithms that enable models to identify their own hallucinations through selective prediction. Together, these contributions aim to establish a new generation of reliable and efficient visual-language systems for safety-critical applications. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria. NSF Award ID: 2536029 | Program: 01002627DB NSF RESEARCH & RELATED ACTIVIT | Principal Investigator: Nuno Vasconcelos | Institution: University of California-San Diego, LA JOLLA, CA | Award Amount: $600,000 View on NSF Award Search: https://www.nsf.gov/awardsearch/show-award/?AWD_ID=2536029 View on Research.gov: https://www.research.gov/awardapi-service/v1/awards/2536029.html
Interested in this grant?
Start a free 7-day trial to get match scores, save grants, and build your application with AI.
Grant Details
$600,000 - $600,000
Not specified
LA JOLLA, CA
View the application link
Start a free 7-day trial to open the original listing and funder website, save this grant, and track its deadline. Cancel anytime.
Start free trialWant to see how well this grant matches your organization?
Get Your Match Score