As militaries drown in disconnected data, Virtualitics is betting that readiness — not the battlefield — is where agentic AI proves whether it can be trusted with high-stakes decisions.
Military readiness depends on knowing, in real time, whether people, equipment, and supplies can actually deploy. That answer typically sits scattered across a dozen disconnected systems, reconciled by staff working from reports that are days or weeks old. Virtualitics is trying to close that gap with Iris, an agentic AI platform built specifically for defense and national security decisions, where a wrong or unexplainable answer carries real operational consequences.
In this interview, Aakash Indurkhya, Chief Product Officer at Virtualitics, explains how Iris layers purpose-built machine learning models with large language model-powered agents to reason across changing data instead of simply retrieving it. He also addresses the central tension in government AI adoption: how a system can rely on foundation models widely criticized as opaque while still providing commanders with a fully traceable, auditable chain of evidence for every recommendation.
Full interview;Â
Virtualitics built Iris around the idea that military readiness, knowing if your people, equipment, and supplies are actually ready, is where AI can make the biggest difference. What does that AI actually look like under the hood when the stakes are a real military mission?
Most enterprise software was built around dashboards and workflows, and then AI was bolted on as an add-on. We built Virtualitics Iris assuming AI would be the user’s partner from day one. That changes the architecture substantially.
Under the hood, there are layers of AI working together. At the base, specialized data engineering, ontology management tools, and machine learning models do the heavy analytical work: scoring equipment and assets, projecting risk, and forecasting impacts. These are not generative models. They are purpose-built models that create a structured foundation of analytics insights.Â
On top of that, LLM-powered agents enable analytic specificity. When a user asks “What are the top readiness risks for units in the patriot battalions deploying in the next 90 days?”, an agent pulls in conversational context, decomposes the question into a plan, calls on those underlying ML tools, writes and executes code to pull and filter the right data across multiple sources, reconciles conflicting information, and synthesizes everything into a response with full source attribution. The user sees where every number came from.
The key difference from traditional analytics is that AI orchestrates the entire workflow rather than just answering a question. It reasons across changing data, cites the evidence it used, and helps navigate decisions rather than just presenting charts. The user is not navigating software. The software is helping them navigate a decision.
A lot of organizations have more data than they know what to do with, but still end up making slow, bad decisions. Where exactly does that break down — and what does Virtualitics fix that a regular analytics tool doesn’t?
Traditional analytics end at “here’s what happened.” Virtualitics goes further: “Here’s what’s happening, why it’s happening, what will likely happen next, what your options are, and the trade-offs between them.”
The problem today is not a lack of data. It is that the data lives across disconnected systems, in too many reports, too many dashboards, too many alerts, and with not nearly enough situational specificity to inform the decision in front of you right now. Even a great analytics system surfaces hundreds of potentially interesting insights. But the person making the decision still has to manually assemble them, determine which are relevant to their specific situation, and reconcile any conflicts. That assembly step is where decisions stall.
Agentic AI changes that because agents do not just retrieve information. They gather evidence across data sources, reconcile conflicting signals, identify the actual constraints, and generate courses of action with explanations. When an officer gets conflicting personnel and equipment data across three different reporting systems, Virtualitics Iris reconciles them and surfaces the real constraint rather than leaving the officer in a chair-swiveled mess of dashboards.
Also Read: Why Most Enterprise AI Agent Pilots Never Reach Production
You just announced a partnership with OpenAI. Government customers can’t compromise on knowing exactly why an AI made a call, but the models powering it are notoriously opaque. How do you square that?
We treat the foundation model as a reasoning engine, not an oracle. Models like OpenAI’s GPT offer extraordinary reasoning capabilities, but mission decisions require far more than that. They require domain contextualization, data provenance, traceability, and guardrails. That is where an application like ours and a foundation model become far more powerful together than either could be on its own.
Concretely, when Virtualitics Iris answers a readiness question, it shows you exactly which data sources it queried, what filters and joins it applied, and links back to the underlying records. Every response carries source attribution. The model handles the reasoning; the application structures the problem, grounds it in real data, and delivers the evidence chain. So the user does not have to take the AI’s word for it. They can trace any claim back to the data and challenge it.
That is the core of how we square it. We do not try to make the model itself transparent. We build an application layer around it that makes every output traceable, auditable, and challengeable. Government customers need to be able to interrogate a recommendation before acting on it. We designed the system to support exactly that.
After seven years of building AI for defense and critical infrastructure, what do most enterprise AI leaders genuinely get wrong about building AI that people actually trust with important decisions?
The biggest misconception is that trust comes after you have built the AI. In reality, trust has to be engineered into the system from the beginning. Too many AI solutions optimize for speed to answer, but in mission-critical environments, the answer alone is never enough.
People need to understand why a recommendation was made, what data it was based on, what assumptions were used, and how confident the system is. They need to be able to interrogate the recommendation, trace it back to the underlying evidence, and challenge it if something does not look right. That is especially true in national security, where decisions have real operational consequences.
We have always made transparency and chain of thought core design principles. The user’s question is not just a prompt; it is the entry point into an entire analytical workflow they can follow, inspect, and redirect. The system shows its work at every step, and the user can spot-check any piece of it. That is what makes the difference between AI people who tolerate it and AI people who actually rely on it when the stakes are high.
Also Read: The Hidden Cost of AI Adoption Without a Plan
You’ve called readiness “the front door for agentic AI in government.” Why readiness specifically, and where does this all go if it actually works at scale?
Readiness is a perfect proving ground for agentic AI because it is one of the most data-rich domains in the federal government, yet remains poorly coordinated across a vast decision landscape that is always changing. Personnel, training, logistics, maintenance, equipment, supply chain, mission priorities, operational risk–all of it lives in disparate systems. As a result, decisions are either slow or under-informed.
That matters beyond just efficiency. Today, the people who keep readiness operations moving often do so through tribal knowledge about how data across these systems should be connected and interpreted. When those people rotate or leave, that institutional knowledge goes with them, jeopardizing the ability to inform readiness decisions the same way. Agents provide both continuity and speed.
Here is a concrete example: AI-projected component risk, downstream maintenance work orders, and supply forecasting could all be synthesized by an agent focused on flight scheduling. That level of coordination currently requires multiple humans across multiple systems, and becomes infeasible at the scale of decisions that have to be made weekly. A traditional algorithm would not work either, because you need to capture and convert situational details that would be too tedious to encode through a traditional software interface. That is exactly what agentic AI is built for.
If this works at scale, a commander could assess the readiness impact of a supply chain disruption and receive a synthesized answer in minutes, with full traceability to the underlying data. That is the trajectory we are on.
If Virtualitics gets this right over the next decade, what actually changes — in how militaries operate, how governments make decisions, how humans and AI divide responsibility in high-stakes situations?
If we get this right, AI becomes part of every decision that impacts the readiness mission. Not as a replacement for human judgment, but as a system that continuously synthesizes information, anticipates risks, evaluates options, and surfaces the best course of action so the decision-maker can act with confidence instead of uncertainty. Done well, that translates directly into faster, better-informed readiness decisions across the force, for our military and our allies.
Today, a combatant command staff assembles readiness pictures across all subordinate units, often working from reports that are days or weeks old. In a decade, that picture should be synthesized in real time, with AI agents monitoring every data feed, flagging changes that matter, and proactively surfacing risks before they become crises. Decisions that currently take weeks of staff work should take hours.
The division of responsibility may not change. Humans will own the judgment, the priorities, and the accountability. What changes is how quickly and confidently they can act, because the information they need to make good decisions is actually assembled, reconciled, and ready when they need it rather than scattered across a dozen systems and someone’s institutional memory.
Also Read: AI-Generated Code Needs a New Testing Model
You studied computer science, spent years deep in the technical side of AI, and now you’re running a product. At what point did you have to stop thinking like an engineer, and did you?
I never stopped thinking like an engineer. I just expanded what I was engineering. Early in my career, success meant building the best algorithm. The shift happened when I started seeing all the barriers that stop people from actually leveraging what those models produce. Sometimes it is a trust issue. Sometimes the analysis is not specific enough to the situation at hand. Sometimes people need help structuring a problem before they can even ask the right question. And sometimes the model does not account for the right constraints, and the user knows it.
That is what I spend my time on now: eliminating those barriers. The technology is only successful if it helps someone make a better decision. That has not made me less technical. It has made me more deliberate about what I choose to build and why.


