Search Contract Opportunities

Artificial Intelligence Performance Evaluation Tool

ID: DAF26TZ06-NV006 • Type: SBIR / STTR Topic • Match:  95%
Opportunity Assistant

Hello! Please let me know your questions about this opportunity. I will answer based on the available opportunity documents.

Please sign-in to link federal registration and award history to assistant. Sign in to upload a capability statement or catalogue for your company

Analyze Opportunity
Loading

Description

TECHNICAL POINT OF CONTACT (TPOC)
Frank Cruz
Jaime Ramirez
PROJECTED CMMC LEVEL REQUIREMENT
Level 2 (Self)
TECHNOLOGY AREAS
Information Systems
|
Sensors
MODERNIZATION PRIORITIES
Trusted AI and Autonomy
KEYWORDS
Artificial Intelligence; AI Performance Evaluation; Trusted AI; Autonomy; Explainability; AI Metrics; AI Validation; AI Systems; Sensor Technologies; Information Systems; Operational Testing; ABMS; 412TW; AFTC; DoD AI Standards; Test and Evaluation.
OBJECTIVE
The objective of this research effort is to design, develop, and demonstrate an Artificial Intelligence (AI) Performance Evaluation Tool that streamlines, standardizes, and enhances the evaluation of AI-enabled systems across various operational conditions. This Phase I STTR effort will focus on exploring innovative methodologies to quantify the performance, reliability, and robustness of AI models, particularly in applications related to sensors and information systems within the test and evaluation missions of the Air Force Test Center (AFTC). To demonstrate early-stage feasibility, the proposed tool must be capable of evaluating AI-enabled sensor suites (e.g., EO/IR, Radar, and Radio Frequency systems) and align with AFTC's primary simulation and digital engineering environments, such as the Advanced Framework for Simulation, Integration, and Modeling (AFSIM), Ansys Test and Evaluation Tool Kit (TETK), or the Joint Simulation Environment (JSE).
By addressing critical gaps in assessing AI robustness under dynamic and contested conditions including real-time inference latency, adversarial resilience, and data drift detection the effort aims to establish early-stage feasibility for an adaptive framework capable of identifying key performance indicators (KPIs), assessing trustworthiness through explainability metrics, and ensuring compliance with evolving Department of Defense (DoD) AI standards, such as those delineated in the DoD Responsible AI Strategy and Implementation Pathway.
This tool will be designed to integrate seamlessly with the current test environment of the 412th Test Wing and support Operationally Focused Advanced Battle Management System (ABMS) initiatives by providing actionable insights that enable faster and more informed decision-making processes. Phase I will emphasize technical innovation, feasibility validation, and initial use-case alignment to ensure the framework's scalability and adaptability for Phase II development and field-level implementation. The alignment of the initial use case must specifically demonstrate how the proposed evaluation tool will ingest AFTC-representative sensor data to ensure the framework's scalability and adaptability for Phase II development. The overarching goal is to accelerate the operational utility and trusted deployment of autonomous and AI-driven systems in support of Air Force modernization priorities while improving mission readiness and capability sustainment.
DESCRIPTION
The proposed research effort aims to address critical gaps in the evaluation and validation of Artificial Intelligence (AI) systems within the Air Force by developing an Artificial Intelligence Performance Evaluation Tool. The tool will provide a standardized, adaptive, and robust framework capable of assessing the performance, reliability, and explainability of AI models, with a targeted application in sensor and information systems used in test and evaluation missions under the 412th Test Wing (412TW) of the Air Force Test Center (AFTC). This effort aligns with the Department of Defense's (DoD) thrust toward advancing Trusted AI and Autonomy and operationally focused Advanced Battle Management System (ABMS) initiatives.
Problem/Unmet Need: Current AI evaluation processes lack standardization and reproducibility across varying operational environments, making it difficult for the AFTC to confidently assess AI models' trustworthiness, reliability, and operational readiness. Furthermore, the escalation in complexity of AI-enabled systems has underscored the need for technologies that can quantitatively measure AI performance while adhering to DoD Responsible AI principles. These challenges prevent the Air Force from leveraging AI advancements at scale, impacting readiness and modernization goals.
Opportunity/Desired Outcome: By creating a modular AI evaluation tool, this effort provides an opportunity to establish a scalable and adaptable solution that aligns with emerging DoD AI standards. The desired outcome is a measurable enhancement in the fidelity, speed, and consistency of AI performance assessments, reducing decision-making cycles and increasing mission readiness. The tool will enable practical implementation of AI validation in secure environments, ensuring compliance with applicable standards and operational conditions. Successful deployment will support the broader modernization priorities tied to mission-critical AI deployment and advanced autonomy.
Approach
The project will begin in Phase I with baseline research and feasibility studies, advancing into a scalable prototype in Phase II for integration and testing in operationally relevant environments. A summary of planned efforts by phase follows:
Phase I: Feasibility and Concept Development (Initial TRL: 2; Target TRL: 4)
Research: Conduct an initial landscape analysis of AI performance metrics, methodologies, and tools with a focus on gaps and challenges specific to AFTC's mission space. Identify and assess algorithms for measuring real-time inference latency, adversarial resilience, and data drift specifically for different sensor models. This research must establish and utilize open-source or synthetically generated surrogate sensor datasets that are representative of AFTC flight test scenarios.
Framework Design: Develop a software architecture for an adaptable AI performance evaluation framework designed to address reliability, robustness, and explainability metrics. Awardees will need to base their design on widely accepted DoD digital engineering and telemetry standards. The architecture must explicitly define the assumed data ingestion pipelines for surrogate sensor data and detail the computational methods that will be used to calculate reliability, robustness, and explainability metric.
Simulation: Perform initial simulation tests to explore early-stage integration of explainability metrics and KPI identification methodologies. This will consist of running dry-runs within a localized software sandbox. The objective of these tests is to demonstrate that the proposed mathematical models can successfully process these inputs, identify Key Performance Indicators (KPIs) such as data drift, and output quantifiable explainability scores.
Deliverables: A feasibility report describing how the tool can practically operate within the AFTC environment, a detailed framework design document, and initial simulation results demonstrating concept validity. This could include Surrogate Data Strategy and Baselining, algorithmic and mathematical validation, architectural alignment and computational overhead, and Phase II scalability and transition pathway.
Phase II: Prototype Development and Validation (Initial TRL: 4; Target TRL: 6)
Prototype Development: Build a functional prototype of the performance evaluation tool with a secure Python-based backend and a user-friendly Graphical User Interface (GUI).
Integration: Test and integrate the tool with existing sensor and information systems in a controlled 412TW test environment to assess real-world performance.
AI Metrics Library: Develop an expandable library of explainability metrics aligned with DoD Trusted AI standards to support diverse mission applications.
Validation: Conduct comprehensive testing in operationally relevant scenarios to ensure reliability, scalability, and compliance with DoD standards.
Deliverables: A fully functional prototype, test reports validating system performance, and a roadmap for field-level deployment.
Conclusion
This effort supports the Air Force's operational imperatives by advancing capabilities for trustworthy AI deployment in mission-critical applications. Upon completion of Phase II, the AI Performance Evaluation Tool is expected to reach TRL 6, with a clearly defined transition pathway toward operational implementation across the 412TW and other DoD components. This innovation is anticipated to significantly enhance the Air Force's ability to validate AI systems, improve decision-making confidence, and maintain technological superiority in contested environments.
PHASE I
Phase I will establish the project's feasibility and lay the foundational groundwork. Key activities will include: Detailed Analysis: Conduct a comprehensive study of current flight test analysis workflows to identify specific bottlenecks, constraints, and user requirements. Architectural Design: Develop a detailed hardware and software architectural design that emphasizes modularity, security, and Python-based extensibility. Core Algorithm Prototyping: Develop and simulate proof-of-concept algorithms for automated metric generation and data processing. Feasibility Report: Deliver a final report summarizing findings, the complete system design, simulation results, and a detailed plan for the Phase II prototype development.
PHASE II
Phase II will focus on the creation and rigorous testing of a fully functional prototype in a controlled environment. Prototype Development: Build the integrated software application, including the graphical user interface (GUI), data analysis engine, and the AI-powered report generation module. Lab-Based Testing: Establish a reference system in a secure, closed-loop environment to conduct extensive testing. Performance Validation: Evaluate the prototype against a comprehensive set of metrics, including: Time Space Position Information (TSPI) processing accuracy. Comparative analysis against legacy methods. Integrity of automated report generation. System security and resilience against simulated threat/victim scenarios. Demonstration: Conclude with a successful end-to-end demonstration of the prototype's capabilities, proving its readiness for transition to a real-world operational testing environment.
PHASE III DUAL USE APPLICATIONS
The 416th Test Squadron at Edwards Air Force Base, which specializes in F-16 flight testing and has a history of integrating advanced systems, will serve as the dedicated transition partner for this effort. Our commercialization strategy is multifaceted: Productization: Refine the successful Phase II prototype into a robust, user-friendly, and cost-effective commercial product. User-Centric Feedback: Conduct extensive operational testing with the 416th Test Squadron and other target users to gather feedback for iterative product improvements. Dual-Use Applications: While the primary customer is the DoW, we will explore dual-use applications for commercial aerospace, avionics testing, and other sectors that require complex data analysis. Infrastructure and Support: Establish comprehensive manufacturing, customer support, and training infrastructure to ensure a smooth and successful transition from a prototype to a fully supported operational tool that will revolutionize flight test analysis.
REFERENCES
Department of Defense (DoD) Responsible Artificial Intelligence Strategy and Implementation Pathway. https://www.ai.mil/responsible_ai.html
National Institute of Standards and Technology (NIST) Artificial Intelligence Risk Management Framework. https://www.nist.gov/ai-risk-management
Air Force Test Center (AFTC) Organizational Overview. https://www.aftc.af.mil/
John M. McQuade, et al. Metrics for Trustworthy AI: Evaluating Reliability, Robustness, and Explainability in Defense Applications. DoD Artificial Intelligence Symposium, 2023.
Executive Order 13960 - Promoting the Use of Trustworthy Artificial Intelligence in the Federal Government. https://www.whitehouse.gov/ai/
Advanced Battle Management System (ABMS) Fact Sheet. U.S. Air Force, Office of the Assistant Secretary of the Air Force for Acquisition, Technology and Logistics. https://www.af.mil/About-Us/Fact-Sheets/Display/Article/1045816/advanced-battle-management-system-abms/
Defense Innovation Unit (DIU) Trusted AI Framework: Operationalizing AI Ethics and Reliability. March 2022. https://www.diu.mil/
Test & Evaluation (T&E) Framework for Autonomy. Presented at T&E Working Group, AFTC, 2022.
Explainable AI (XAI) Principles for DoD Applications. Defense Advanced Research Projects Agency (DARPA) XAI Program Summary, 2021. https://www.darpa.mil/program/explainable-artificial-intelligence
Office of the Under Secretary of Defense for Research and Engineering. Key Performance Indicators for AI and Autonomy Testing. Report No. OUSD(A&S)-21-0045, 2023.
QUESTIONS & ANSWERS
NOTE:
To ask a question, you must
log in or create an account
for the DSIP.

Overview

Response Deadline
Due in 40 Days
Posted
Open
Set Aside
Small Business (SBA)
Place of Performance
Not Provided
Source
Alt Source

Program
STTR Phase I
Structure
Contract
Phase Detail
Phase I: Establish the technical merit, feasibility, and commercial potential of the proposed R/R&D efforts and determine the quality of performance of the small business awardee organization.
Duration
1 Year
Size Limit
500 Employees
Eligibility Note
Requires partnership between small businesses and nonprofit research institution
On 9/2/26 Department of the Air Force issued SBIR / STTR Topic DAF26TZ06-NV006 for Artificial Intelligence Performance Evaluation Tool due 10/21/26.

Documents

Posted documents for SBIR / STTR Topic DAF26TZ06-NV006

Opportunity Assistant


Analyze Opportunity

Contract Awards

Prime contracts awarded through SBIR / STTR Topic DAF26TZ06-NV006

Incumbent or Similar Awards

Potential Bidders and Partners

Awardees that have won contracts similar to SBIR / STTR Topic DAF26TZ06-NV006

Similar Active Opportunities

Open contract opportunities similar to SBIR / STTR Topic DAF26TZ06-NV006

Experts for Artificial Intelligence Performance Evaluation Tool

Recommended subject matter experts available for hire