Industry — Insurance

AI for insurance, across underwriting and claims.

InTheCloud builds and ships AI for insurers — underwriting and submission triage copilots, claims and FNOL automation, and document intelligence — evaluated against your own cases and deployed inside your cloud under SOC 2 and model governance controls.

We deliver AI implementation and advisory for P&C, life, and specialty carriers, MGAs and brokers. Engagements start with a scored use case and a target metric, reach an evaluated prototype in 4 to 8 weeks, and land in production integrated with Guidewire, Duck Creek or your bespoke platforms — with the audit trail a model risk review expects.

Lines of business
P&C, life, specialty and reinsurance — personal and commercial lines, including MGA and broker workflows.
Platforms we integrate
Guidewire, Duck Creek, bespoke policy and claims systems, document stores, and data warehouses already in your estate.
Controls
SOC 2 Type II and ISO 27001 practice, PII redaction before inference, prompt and tool-call audit logging, documented model governance.
Typical timeline
4 to 8 weeks to an evaluated proof of value; 3 to 6 months to a production system integrated with policy and claims platforms.

How we deliver

01

Underwriting & claims framing

We work with underwriting, claims, and IT leaders to pick AI use cases that move loss ratio, cycle time, or expense — with a clean regulatory story.

02

Compliant prototype

We ship a working agent or model against your submission, policy, and claims data with evaluation against accuracy, hallucination, and policy guardrails.

03

Production engineering

We integrate with Guidewire, Duck Creek, document stores, and downstream systems, and pass security and compliance review.

04

Governance & handover

We leave behind MLOps, model governance, and trained engineering teams aligned to your AI risk framework.

Implementation patterns we ship

P1

Submission triage and appetite matching

Incoming submissions are parsed from email, PDF and ACORD forms, matched against written appetite rules, and routed with a confidence score. Out-of-appetite risks are declined early with a reason underwriters can read and override.

P2

Underwriting copilot with citation-grade grounding

Retrieval over policy wordings, loss runs and guidelines, with every generated statement carrying a link back to the source page. Underwriters keep the decision; the model removes the reading.

P3

FNOL and claims document intelligence

First notice of loss intake, coverage checks against the bound policy, and extraction from medical, repair and legal documents, with structured output written straight into the claims platform.

P4

Fraud and leakage signal review

Models surface anomaly and inconsistency signals across the claim file for a human investigator, with the evidence trail attached. No automated denials — the pattern is triage, not judgment.

P5

Eval harness and drift monitoring

Golden datasets from your real cases, rubric graders, and adversarial examples run on every prompt change, with regression gates in the pipeline and drift alerts on live traffic.

P6

Regulator-ready audit trail

Every inference records inputs, retrieved sources, tool calls, model and prompt version, and the human decision that followed — the record a market conduct or model risk review asks for.

P7

Subrogation and recovery identification

Claim files, police reports and repair estimates are read for third-party liability indicators that the adjuster would otherwise find late or not at all, surfaced as a ranked worklist with the supporting passages attached.

P8

Policy wording and endorsement comparison

Structured diffing of manuscript wordings, endorsements and schedules against your standard forms, flagging deviations by clause with the coverage consequence written in plain language for the underwriter.

P9

Broker and policyholder service copilot

Contact-centre and broker-service assistants grounded in the bound policy, the claim history and the state-specific rules, drafting responses for licensed staff to approve rather than sending unattended.

P10

MCP tool layer over policy and claims systems

A least-privilege tool layer that lets models read Guidewire, Duck Creek or bespoke systems under the individual user's entitlements, with idempotent writes and human confirmation on anything irreversible.

Delivery highlights from adjacent engagements

Data — governed customer data platform

Consolidated fragmented customer data into one governed platform with lineage and access controls — the same retrieval substrate an underwriting or claims copilot has to read from before it can be trusted.

Read: Data — governed customer data platform

Evaluation harness for regulated workloads

How we build the offline evaluation suite — golden datasets, rubric graders calibrated against human reviewers, CI release gates and drift monitoring — that lets a carrier change prompts and models without a committee each time.

Read: Evaluation harness for regulated workloads

Least-privilege model access to systems of record

Our implementation guide for production MCP servers: token pass-through identity, tool scoping, idempotent writes and adversarial testing — the pattern behind copilots that read policy and claims systems safely.

Read: Least-privilege model access to systems of record

LLMs in production, gated by evals

Applied large language models to a live commercial workflow with an evaluation harness gating every prompt change before release, rather than shipping on demo quality.

Read: LLMs in production, gated by evals

Bring us the AI initiative you're trying to get into production.

Scope an insurance AI engagement

What happens next

  • 30-minute builder-led call
  • No generic sales pitch
  • Architecture, feasibility and constraints
  • A recommended next step

Where does AI actually pay for itself in an insurance business?

Three places, consistently. Submission and claims intake, where documents are read and re-keyed by people today. Decision support, where an underwriter or adjuster spends more time locating information than judging it. And service, where routine broker and policyholder questions absorb licensed staff.

The pattern that fails is the one aimed at the decision itself. Carriers that automate the reading and retrieval around a decision, and leave the decision with the human, get results that survive both a model risk review and a bad quarter.

We insist on a measurable target before building: cycle time on a named workflow, expense per claim, quote turnaround, or straight-through processing rate. If nobody will commit to a number, the use case is not ready.

What does a regulator-ready insurance AI system look like?

It is auditable at the level of a single decision. For any given quote or claim, you can produce the inputs, the documents retrieved, the tool calls made, the model and prompt version in force, and the human action that followed.

It is bounded. The model has least-privilege access through a defined tool layer, writes are idempotent and reversible, and irreversible actions require human approval. PII is redacted before inference and inference stays inside your cloud account.

And it is measured continuously, not once at launch: an evaluation suite in the pipeline, drift alerts on live traffic, and documented governance that maps to your existing model risk framework rather than inventing a parallel one.

Build, buy, or extend the insurtech platform you already have?

Buy where the workflow is generic and the vendor's data advantage is real — standard document extraction and off-the-shelf contact-center automation are usually not worth building.

Build where the workflow encodes your appetite, your wordings, and your claims philosophy. That is the part a vendor cannot ship you, and it is where the margin difference between carriers actually lives.

In most engagements the answer is a thin layer of your own over platforms you already run. We are happy to tell you the buy answer when it is the right one — we would rather scope the work correctly than sell a build.

How do NAIC and state model-governance expectations shape the build?

The NAIC model bulletin on AI, now adopted in most states, asks carriers for a written AI governance program: an inventory of AI systems, documented risk assessment per use case, testing for unfair discrimination, third-party model oversight, and evidence that a human is accountable for outcomes. None of that requires new technology; it requires artifacts that most pilots never produce.

So we produce them as we build. Every system we ship carries a one-page system description, the evaluation dataset and grader definitions, bias and disparate-impact testing on the protected characteristics your legal team identifies, a vendor and model dependency list, and a named business owner. The pack is refreshed each quarter rather than written once for launch.

For third-party models the same rules apply to your provider. Deploying Claude or Gemini through Bedrock or Vertex inside your own account is what makes that story straightforward: data residency, retention and sub-processor questions get standard answers your cloud agreement already covers.

What does the data and document layer have to look like first?

Carrier document estates are the real constraint, not model quality. Submissions arrive as email bodies, scanned PDFs, ACORD forms, spreadsheets and broker portals. Loss runs are inconsistent across brokers. Policy wordings live in a document store that nobody has indexed since the last migration. A copilot built over that estate is only as good as the retrieval underneath it.

The first weeks of most insurance engagements are therefore ingestion and retrieval work: OCR with layout preservation for scanned documents, table extraction that survives merged cells, document classification and de-duplication, chunking that respects clause boundaries in wordings, and metadata that carries policy number, effective dates and line of business so retrieval can be filtered before it is ranked.

We measure that layer on its own before any generation is judged: retrieval hit rate on a labeled set of questions, extraction accuracy per field, and coverage of the document types actually present. If retrieval is weak, no prompt fixes it — and knowing that early is what keeps a program honest.

How do you avoid an unfair-discrimination problem in underwriting or claims?

By keeping the model away from the decision and testing anyway. In every underwriting or claims system we ship, the model retrieves, summarizes, extracts and drafts; the rating, pricing, coverage and denial decisions stay in your existing rules and with your licensed staff. That boundary is the strongest single control available, and it is architectural rather than aspirational.

Then we test. Proxy variables are the risk — ZIP code, occupation, vehicle, named entities in free text can all carry protected characteristics indirectly. Evaluation includes paired testing on matched cases that differ only on a sensitive attribute, and we look at outcome distribution across segments, not just aggregate accuracy.

And we log at the decision level, so that if a market conduct examination asks why a specific claim was handled the way it was, the answer is a record rather than a reconstruction.

What actually goes wrong in carrier AI programs?

Four failures repeat. A pilot on synthetic or sampled data that never faces the real document distribution, and collapses on contact with production. A use case with no committed underwriting or claims owner, which produces a system nobody adopts. Security and data-access work started after the build instead of alongside it, which turns a 30-day pilot into a 90-day one. And no evaluation harness, which means nobody can change a prompt after launch without a committee.

The counter-measures are unglamorous and they are why we sequence engagements the way we do: real data in week one, a named business owner before kick-off, the security workstream opened in parallel, and an evaluation suite delivered with the first working prototype rather than after it.

We would rather return a well-evidenced no on a use case in six weeks than deliver a demo that funds an eighteen-month program on the wrong problem.

Frequently asked questions

What insurance AI workloads do you ship?

Underwriting copilots, submission triage, claims FNOL automation, document and policy intelligence, fraud signal review, broker and contact-center copilots, and actuarial and reserving support tooling.

What do AI advisory services for insurance actually include?

Use-case scoring against loss ratio, cycle time and expense; data and document readiness assessment; model and platform selection; a regulatory control story agreed with compliance; and a costed delivery roadmap. We advise and then build — the advice is written by the engineers who will ship it.

How do you handle policy, PII, and regulator scrutiny?

We operate under SOC 2 Type II and ISO 27001, redact PII before egress, audit-log every prompt and tool call, and document model governance for state and federal regulators.

Do you work with P&C, life, and specialty carriers?

Yes. Our engineers have shipped systems for P&C, life, specialty, and reinsurance carriers, and we integrate with Guidewire, Duck Creek, and bespoke policy and claims platforms.

How does this fit with our existing insurtech stack?

We build around the policy administration, claims and document systems you already run rather than replacing them. Integration is through APIs and event streams, with an MCP or service layer that gives models least-privilege access to the systems of record.

Which AI models do you deploy for insurance?

Anthropic Claude, OpenAI GPT, and Google Gemini, deployed via AWS Bedrock, Azure OpenAI, or Google Vertex inside customer VPCs — chosen per workload on cost, latency, and evaluation results.

How long before an insurance AI system is in production?

A scoped proof of value against your own submissions or claims typically runs 4 to 8 weeks. Production deployment, including security review and integration with policy and claims platforms, usually lands within 3 to 6 months.

Who from our side needs to be involved?

An underwriting or claims owner who can define a good outcome, a data or platform engineer for access, and someone from risk or compliance early rather than at the end. Typical carrier-side commitment is a few hours a week, not a full-time team.

Scope an insurance AI engagement

Related reading

AI implementation services

End-to-end AI delivery for enterprise teams.

AI for financial services

Production AI across banking, insurance and fintech.

Agentic AI consulting

Multi-step agents, tool use and bounded autonomy.

What AI implementation costs

How engagements are scoped and priced.

Ready to get back to building?

Tell us about the engagement. We typically respond within one business day with a named Builder who can talk substance — not a generic sales pitch.

Start the conversation

Prefer email? info@inthe.cloud

What happens next

  • 30-minute builder-led call
  • No generic sales pitch
  • Architecture, feasibility and constraints
  • A recommended next step