V

VoxNova

Explored

ABC 2026 · June — August 2026

Explored AI reliability across Insurance & claims processing, Customer Service, Geospatial vision, and Warehouse logistics — mapping where AI breaks and how to make it trustworthy.

Team VoxNova

Team Members

Anirudh Nagarajan

Anirudh Nagarajan

Suma P

Suma P

Prajanth B

Prajanth B

P

Pragna Podamala

Sharvanee Naru

Sharvanee Naru

S

Srishti Shetty

Field Notes

VoxNova explored AI reliability across four verticals where AI failure has significant real-world consequences: 1. Insurance & Claims Processing - Investigated where AI models fail during claims adjudication: edge cases, ambiguous policy language, and adversarial inputs that cause incorrect approvals or denials. Explored how to make AI decisions auditable and contestable. 2. Customer Service AI - Mapped failure modes in LLM-based customer service agents: hallucinations, context loss across long conversations, inappropriate tone escalation, and failure to escalate to human agents at the right moment. 3. Geospatial Vision - Explored reliability challenges in computer vision models applied to satellite and aerial imagery: domain shift between training data and deployment geography, seasonal variation, and cloud cover artifacts. 4. Warehouse Logistics - Investigated AI reliability in pick-and-place robotics and inventory forecasting: how models degrade when SKU mix changes, and how to design fallback systems when AI confidence drops below threshold. Core thesis: AI reliability is not a post-deployment concern - it must be designed in from the start. The team explored frameworks for measuring, monitoring, and improving AI reliability across these high-stakes verticals. Verification Layer for Customer Service AI Agents Domain: Customer Service × AI Background and Existing Processes Customer service organizations increasingly utilize AI agents to handle live interaction functions like ticket resolutions, refunds, account updates, and policy inquiries. These AI agents are provided with considerable independence to perform actions rather than just recommend a human agent the right answer. Identified Problem and Stakeholders Most of the current AI agents operate without any verification layer, which means the response/action is not verified by an independent entity before being sent to the customer/system of record. The affected parties here include customers, customer service teams, approval teams, and the organization using the AI agent. Core Pain Points The response of agents goes directly to the customer or the record system without any stage of validation independently. Consequently, incorrect actions are detected after the fact through complaints from customers or audits. The autonomy obtained by agentic AI systems has been growing quicker than the available verification mechanism for their decisions. The decisions that are not verified in interactions that are regulated or take place in high trust generate audit, compliance, and reputation risks for the company. Results of Research The research that was conducted on the current customer service AI deployment indicates that with the increasing autonomy level of agents, the absence of an independent verification stage indicates a growing operational and trust gap instead of a minor implementation point. Problem Statement Reformulation It is crucial to introduce an independent verification solution that will help check the outputs of the customer service AI agent before being delivered to the end-user. Preventing AI Hallucinations in Customer Service Area of Interest: AI and Customer Service Background and Existing Processes AI support agents and chatbots exhibit hallucination; that is to say, the output they generate may be factually untrue, although they may seem credible (wrong terms of policies, wrong status of orders and accounts, etc.). Identified Problem and Stakeholders Most currently available strategies of hallucination avoidance deal with the AI output verification, while fewer of them verify the assumptions on which the original customer question was based on. This leads to loopholes allowing for providing wrong-premise-based responses. Participants affected are customers, customer service agents using AI solutions, and firms providing the services. Core Pain Points There are times when responses given seem quite understandable and so convincing that knowing they are false may only be possible for a few. Most current solutions have safeguards in place to check the output generated but do nothing to check whether the query posed is false. Giving wrong information about policy, billing or order status will harm the reputation of AI systems in customer service. If a customer is given an inaccurate answer, they will need to contact customer service again, which will negate the efficiency of the system. Results of Research The research revealed that this situation is more complicated than simple errors; this study pointed out a particular gap in the validation of the premise stage. Problem Statement Reformulation There is a need in the solution which checks whether the premise of queries is correct prior to generating answers, besides verifying the output afterward, in order minimize the false or sycophantic hallucinations of AI-based customer support. EvidenceIQ — Complex Liability Claims Investigation Field of Activity: Insurance & Claims Background and Existing Processes Simple claims in the insurance sector have become more digital, but complicated and evidence-heavy disputes such as subrogation cases, multi-party liability cases, and product liability cases still need weeks or months of manual investigations. Such cases require evidence in the form of police investigations, telematics data, engineering reports, medical records, photographs, dash-cam, or drone footage, which are managed by different companies that do not have any common information system. Identified Problem and Stakeholders Almost every document needs to be opened to get the necessary information, compared with others, introductory work is done with engineers as well, which means that a senior adjuster or claims expert is working as a human integration layer. Core Pain Points All cases being disputed take a longer time than those not disputed as the former takes six months or longer. Evidence from third parties like different medical records takes up to 60 or more days to arrive especially in cases where many medical providers of records are involved. Different analyses, tests and opinions from experts take longer to arrive. In some states, government authorities had to put a limit to the duration of the cases like arbitration and even cases before the court. Claim generation is done electronically but the human factor is still needed. Results of Research Research conducted related to processes of claims and subrogation has shown that the delay in processing is caused not because of one step being slow, but because of how evidence is shared between different centers. Problem Statement Reformulation In order to speed up investigations of complex closed cases, the first thing is to improve the process of sharing information among those involved in the process. Auto-Extraction of Topographic Features from Satellite Imagery Industry: Defense x Geospatial AI Background and Existing Processes The problem in question is a case from the defense sector (iDEX DISC 10, Indian Army, Directorate General of Military Operations), which requires the exploitation of artificial intelligence/ machine learning (AI/ML) for the automatic extraction of man-made features (building footprints, roads, railroad tracks) from satellite imagery to enable the quick and precise generation of maps in military command centers. Identified Problem and Stakeholders Currently, the photo interpreters extract these features manually from satellite imagery using a time-consuming process that does not allow keeping up with rapid changes taking place on the ground (for example, due to construction activity or road and railway extension). The relevant interest groups are photo interpreters, map-making teams and decision-makers who use the maps for operational planning. Core Pain Points There are issues associated with slow manual interpretation of cartographic data which is ineffective as it does not keep up with the evolving human-created environment. Maps become outdated quicker than the period it takes for the manual survey to carry out the recent building or infrastructure changes. Mapping is recognized as a bottleneck connection between the receiving of images and their use as valuable information by Command and Control offices. The fact that the data on the geospatial situation in India is collected from various organizations with different standards makes the mapping pipeline harder. Lastly, qualified experts who interpret photos are scarce resources which should be lowered through automation. The software has to work in conditions requiring clearance of security and consideration of some factors affecting the possibility of image analysis. Results of Research The research shows that this is an issue of decision-making processing time rather than a computer vision one since the brief states that reducing the time taken to make a map is the main purpose of the study. Problem Statement Reformulation There is a necessity for the automatic analysis of geographic data gathered by satellites in order to decrease the time needed for command-level decision-making Screen Dependency in Warehouse Picking Operations Area of Study: Warehouse Operations plus HCI plus AI Background and Existing Processes Warehouses and fulfilment centres fulfil a high volume of orders by means of manual picking procedures. As a result, these logistics processes rely heavily upon workers with handheld scanners that deliver picking instructions and confirm locations and completion of tasks. Identified Problem and Stakeholders A field view of the picking functions made evident that workers were often breaking their current physical tasks, for example, walking or picking. The particular issue that was seen during observation of the operations was that in a single pick of four items, a worker checked the scanner, walked for some time, checked the scanner again, confirmed the location, and checked the number of good picked after which the task was completed. The stakeholders on the issue include pickers, packers, supervisors plus the firm itself with different aims of the operations. While pickers are interested in finishing their tasks in the fastest way, packers need correct packing. As for the supervisors, their target is to achieve the highest productivity level; the ultimate goal of the firm is to minimise operational costs. Core Pain Points Workers interrupt physical picking activities to check handheld devices and receive instructions. All interruptions increase completion time, so when multiplied by the number of interruptions over time, it results in massive productivity losses. Available options – handheld scanners, printed lists, warehouse management systems, and supervisor assistance – enable information to be conveyed, but still require the worker to step away from his work and check the screen or consult with the supervisor. Similar products and stacked items make it easier to confuse the items, resulting in false picking and returns, as well as decreased satisfaction of customers. Results of Research Field observations and analysis of workflows proved that the widely accepted opinion that the speed of workers walking influences productivity levels is not entirely accurate. The actual flow analysis has indicated that the factor constraining throughput is not how fast the workers move physically, but how they interact with their devices. This is especially important at the background of intensity increase in warehouse automation, faster introduction of AI for business support, rising labor shortage, and wider use of sturdy wearable devices. Problem Statement Reformulation Exploring techniques for hands-free communication and interaction in real-time for the purpose of maximizing the speed at which picking and picking confirmations are carried on, as opposed to the currently used handheld scanners methods of operation.

Notes & Reflections

Life is short, live it!

Anirudh Nagarajan