CERLAN Healthcare

Find the clinical failures your benchmarks may miss.

Independent expert evaluation for a defined healthcare AI workflow. We examine consequential failure cases before a release, pilot, or product expansion.

Request a 20-minute conversation

A focused evaluation scoped to your product, intended use, and next decision.

The evaluation gap

A passing benchmark can leave important questions open.

Healthcare AI teams need to understand how their product responds when context is incomplete, instructions conflict, or a plausible answer carries clinical consequences. CERLAN designs focused challenges around the workflow you are preparing to deploy and examines the failures that emerge.

The launch offering

Clinical AI Failure Audit

An independent, fixed-scope evaluation of a defined healthcare AI workflow.

What it does

Test the system against its intended use.

Together, we set the evaluation boundaries, test a frozen system version against challenging scenarios, and review the findings for their clinical significance.

When teams use it

Approaching a model release, health-system pilot, new agent deployment, specialty expansion, major update, or clinical-safety review.

Patient-facing agentsClinical copilotsMedical LLM appsVoice agentsWorkflow AI

What we evaluate

Questions that matter in use.

01 / CONTENT

Clinical content

Are important facts missing, misstated, or presented with unwarranted certainty?

02 / REASONING

Reasoning and recommendations

Does the response account for relevant risks, limitations, and escalation needs?

03 / BEHAVIOUR

Workflow behaviour

Does the system stay within its intended role when a request is ambiguous or outside scope?

04 / PATTERNS

Consistency

Do related cases reveal recurring failure patterns?

Cases and review criteria are selected for the specific product. The audit does not produce a universal safety score.

How it works

From intended use to reusable tests.

01Scope

Define the workflow, users, intended use, and outcomes to examine.

02Challenge

Build product-specific cases that probe clinically consequential behaviour.

03Review

Run a frozen system version and obtain independent ratings.

04Examine disagreement

Retain differences in reviewer judgment for analysis.

05Adjudicate

Escalate serious or disputed cases to expertise matched to the question.

06Prioritize

Identify failure patterns and practical remediation priorities.

07Retest

Turn selected cases into a private regression suite for the client.

What clients receive

Findings your team can act on.

The agreed scope determines the final package.

You can trace a finding back to the case, the ratings, and the reasoning behind its priority.
  • Row-level evaluation dataset
  • Severity and failure taxonomies
  • Reviewer agreement analysis
  • Adjudication log
  • Prioritized remediation findings
  • Executive findings report and discussion
  • Private regression suite

Judgment and trust

A review process built around the question.

Expertise matched to the task

Clinical evaluation calls for different judgment at different stages. Depending on the workflow, reviews may involve physicians, residents, nurses, pharmacists, clinical researchers, coders, or other relevant specialists.

Review criteria are defined in advance. Independent ratings are retained, disagreement is examined, and consequential or disputed cases can be escalated for adjudication. No single reviewer is automatically treated as ground truth.

Set data boundaries before work begins

CERLAN starts with synthetic, public, client-generated, or appropriately de-identified material wherever possible. Before an engagement, we agree on the data to be used, who needs access, how findings will be shared, and the retention and deletion terms.

Please do not send identifiable patient information through the website or in an initial inquiry.

About CERLAN

Independent expert evaluation for high-stakes AI.

CERLAN is building independent expert evaluation for high-stakes AI, beginning with healthcare. Its first offering helps healthcare AI teams investigate clinically consequential failures in a specific product workflow and carry the resulting cases into future testing.

Founder

Pulkit Kumar

Pulkit holds an MSc in Surgery and has worked across medical research, clinical and global-health research coordination, and AI evaluation. He founded CERLAN to bring structured, independent judgment to healthcare AI evaluation.

Start a conversation

Preparing a release or pilot?

Tell us what your system does, what decision is coming up, and what you most need to learn from an independent evaluation.

Request a 20-minute conversation

Email [email protected]. Please leave patient information out of your first message.