DataFramer
Financial Services
Insurance
Tech
Connect AI Quality to Business Outcomes
Discover, Diagnose & Fix AI Failures
Scale Expert Review & Ground Truth
Calibrate Judges & Guardrails
Business Intelligence → Connect user signals to outcomes
Quality Intelligence → Find failures, causes, and fixes
Human Intelligence → Put humans in the loop
Judge Intelligence → Align judges with expert judgment
Eval Intelligence → Turn failures into regression coverage
Pricing Blog Research About
Log in Sign up
Financial Services Insurance Tech
Connect AI Quality to Outcomes Discover & Diagnose AI Failures Scale Expert Review & Ground Truth Calibrate Judges & Guardrails
See the complete journey Discover, fix & track accuracy problems Put humans in the loop Build & calibrate judges Evaluate & generate
Pricing Blog Research About
Log in ↗ Sign up
Published Research

Research from the
DataFramer team

Our work on AI evaluation, failure detection, expert review, benchmarks, and LLM reliability - published at peer-reviewed venues and on arXiv.

AAAI · PMLR 2026

All Required, In Order: Phase-Level Evaluation for AI–Human Dialogue in Healthcare and Beyond

Introduces OIP-SCE, an evaluation framework that assesses whether conversational AI systems meet all necessary clinical requirements in the proper sequence - making AI dialogue systems more compliant with healthcare workflows and auditable for clinical review.

Kulkarni · Lyzhov · Chaitanya · Joshi Read paper
arXiv · 2025

INSURE-Dial: A Phase-Aware Conversational Dataset & Benchmark for Compliance Verification and Phase Detection

A benchmark dataset of real and synthetic insurance benefit verification calls annotated for compliance auditing - enabling evaluation of voice agents' ability to detect call phases and verify procedural and informational compliance.

Kulkarni · Lyzhov · Joshi · Chaitanya Read paper
arXiv · 2025

HalluciNot: Hallucination Detection Through Context and Common Knowledge Verification

Introduces HDM-2, a hallucination detection system that identifies inaccurate outputs from large language models by verifying responses against both provided context and general knowledge facts.

Paudel · Lyzhov · Joshi · Anand Read paper

©2026 DataFramer. ©2026 AIMon Labs, Inc. All rights reserved.

About Us Privacy Policy SOC 2 Type II Certified HIPAA Compliant

Get in Touch

Please use your work email address.

Thank you for reaching out. We will be in touch shortly! Please enter a valid email address.