Jev Model in Insurance: Agentic AI Use Cases for Underwriting & Claims
As agentic AI in insurance evolves, specialized models like Jev by TypeSafe1 are replacing general-purpose LLMs for specific tasks. Agentic AI in insurance relies on fast, deterministic routing, and Jev is a decision model built exactly for this. It can be used in multiple insurance domains where classification, prioritization, and detection play a role. The model was released on September 15, 2026, and is very different from large language models. Jev does not chat, does not use chain-of-thought reasoning, and does not generate any text. It decides on text you provide, against answers you define in advance. As insurers implement agentic AI use cases across their operations, they should evaluate where Jev fits into the systems they already run, so that a workflow can take a predetermined output at a much lower cost.
Table of contents
Open Table of contents
- How Agentic AI Uses Jev’s Three Primitives in Insurance
- Conclusion
- FAQ
- Who built Jev?
- What is Jev?
- What is the state?
- What does each primitive return?
- What Are the Fundamental Differences Between Jev and LLMs in Insurance Workflows?
- What are the top agentic AI use cases for Jev in insurance?
- Which operational risk controls should be implemented when delegating insurance cases to Jev?
- Is Jev good in all languages?
- How to ensure auditability and reduce the general risk of using Jev in insurance?
- Sources
How Agentic AI Uses Jev’s Three Primitives in Insurance
Jev does not return prose. It returns three primitives:
- Choice
- Score
- Noul
A call has three fields: state, model, and questions. The state is the text Jev judges. It can be a string, a JSON object, or an array of text.
The questions are a separate field. Below, I provide one concrete example for each primitive, and all requests are valid calls directly within TypeSafe’s
playground using Jev version 1.13 (typesafe/jev-1.13). The numbers are captured from the live runs and could move a little from call to call.
- Choice is a commercial-property occupancy.
- Score is the order in which a claims team should touch a first notice.
- Noul is one underwriting exclusion.
Choice: Categorizing Commercial Property Occupancy
Choice returns the winning option, a probability for every option, and a confidence. Confidence indicates how concentrated that distribution is.
This example is an occupancy question on a commercial-property submission. The state is the operations narrative only. The options are a short
list a carrier could rate this building against: the shop, the offices, and five other trades, plus other and not_stated, so nine in total.
A list this short is enough to show the fundamental concept, but a full tariff book still has to stay inside 255 options, and both escape options
stay on the list.
Jev Request and Resulting Response
The following input provides sample wording from a broker or applicant email submission:
Input
State:
{
"operations_narrative": "Property submission for Meridian Stores plc, a clothing chain.
The risk address is one building at 14 King Street, and the broker has left the occupancy box blank.
The ground floor opens off the high street. It has window displays, a public entrance, tills,
fitting rooms, and sales staff serving customers during ordinary trading hours. The first floor is
reached by a public stair and is a further sales floor, with the same garments on rails. The second
floor and the third floor have no public entrance. Staff reach them from a side door with a pass.
Those floors are fitted as the chain's regional office: desks for buyers and merchandisers, a finance team,
human resources, an e-commerce group, meeting rooms, and a boardroom. The people on those floors do not serve
customers. The covering note says only that the building contains the shop floors and the regional offices,
and it does not say which use should be coded. No manufacturing, no printing works, no food production,
and no warehouse serving other branches are described. The goods in the building are the shop's own
merchandise on the sales floors."
}
Questions:
{
"building_occupancy": {
"type": "choice",
"instructions": "Which occupancy class best fits this building?",
"criteria": {
"retail": "A shop selling goods to the public, including the sales floors of a chain.",
"offices": "Offices for administration, buying, finance, or management, including office floors in the same building as a shop.",
"storage_of_finished_goods": "A warehouse holding goods for delivery, with no sales floor.",
"paper_and_cardboard_converting": "Making packaging from paper or board.",
"woodworking": "Sawing or machining wood.",
"printing_with_flammable_solvents": "Printing with solvent-based inks.",
"food_processing": "A business preparing food or drink.",
"other": "A trade that is none of the classes above.",
"not_stated": "No use of the building is described."
}
}
}
The questions carry the instructions and the criteria. The state is a separate field.
Response

Response:
Retail takes 88 percent, and offices keeps 12 percent. The other seven classes, including other and not_stated, are at zero,
which is why the two figures add to 100 percent. The question asks for one class for the building. The sales floors earn it: a clothing
chain, a high-street entrance, tills, fitting rooms, and merchandise open to the public. The upper floors are why offices does not disappear.
Buyers, finance, human resources, and a boardroom sit there, reached by a staff pass, and that class was written so that office floors in the same
building as a shop still count. The blank occupancy box does not become not_stated because the note itself describes both uses. Floor area, construction,
and the rate are not in this question.
Choice: “retail”
Confidence is 86 percent and is a separate field from the class probabilities.
JSON response from TypeSafe UI:
{
"model": "jev-1.13.0",
"answers": {
"building_occupancy": {
"type": "choice",
"choice": "retail",
"confidence": 0.86,
"probabilities": {
"printing_with_flammable_solvents": 0,
"other": 0,
"offices": 0.12,
"not_stated": 0,
"storage_of_finished_goods": 0,
"food_processing": 0,
"retail": 0.88,
"paper_and_cardboard_converting": 0,
"woodworking": 0
},
"stats": {}
}
},
"usage": {
"input_tokens": 739,
"output_tokens": 107
},
"request_id": "playground_1e33fe768144c234a6183ffdda68c62cc6e",
"evaluation_time_ms": 74.38133600226138
}
Score: Triage for First Notice of Loss (FNOL)
For the primitive Score, I use a first notice of loss (FNOL) on a building owner’s property policy. The owner reported a claim to the insurer, and now the question is how quickly a claims handler needs to pick up the file and work on it. This is effectively a claims triage case. For the sake of simplicity, the example has only one dimension and touches neither reserving, a peril code, nor a choice between building coverage and a liability claim from the tenants.
Score returns a score, a legend that echoes your own defined levels, a probability for each level, and a confidence. The score is the probability-weighted
mean of the level indexes, which start at 0. The levels must be in ascending order, from least to most severe, and each level must describe a concrete situation.
A label such as “medium severity” gives the model nothing to match, so each level has to name a concrete situation.
Jev Request and Resulting Response
Input
State:
{
"loss_notice": "Unread overnight email, forwarded by the broker at 06:40.
The insured owns Hale Court, a three-storey multi-tenant office building.
The sender is the managing agent. No handler has opened the file.
From: Priya Shah, Hale Court Management
To: the broker's claims desk
Subject: Ceiling leak at Hale Court
I am writing from the lobby before the building's normal hours.
Just after five the second-floor tenant rang the out-of-hours number.
A joint in the ceiling void over their comms room has opened. Their night person
has switched the cabinets off and taken the night shift out of that room. They
are standing in the lobby with the door propped so the engineer can see in.
The tenant on the floor underneath found water on the ceiling tiles above the open-plan desks.
That tenant has sent the early staff home. The desks under those tiles are empty.
They have laid plastic sheets across the screens and the printers, and someone from
their side is waiting in the car park with the keys. Our engineer was in the plant room
for the morning round and closed the valve on the riser that feeds that void. He says
the handle is fully over and the pipe on his side of the valve is quiet. He also says
drops are still falling from the ceiling of the floor below onto the plastic sheets.
I have not been on either floor. A plumber has been called. I will write again when
the plumber has been in the void."
}
Questions:
{
"queue": {
"type": "score",
"instructions": "How quickly does a handler need to take this first notice, based only on what the notice says is happening?",
"criteria": [
"The event is over. No occupant is out of use, and no injury is mentioned.",
"Damage is at the insured premises only. The insured can keep operating. No injury is mentioned.",
"Another occupant is out of use, or the insured's operations have stopped. The narrative says the source is under control. No injury is mentioned.",
"Water, fire, or collapse is still in progress, or the narrative names an injury."
]
}
The questions carry the instructions and the criteria. The state is a separate field.
Response

Response:
Level 3 takes 83 percent. That is the level for water still in progress. Level 2 keeps 17 percent, the neighboring level, where other occupants are out of their
rooms and the notice describes the source as held. Levels 0 and 1 are at zero, so the two figures add to 100 percent. The returned score is 2.82. It is the
probability-weighted mean of the indexes. On the probabilities as printed, 0.17 × 2 + 0.83 × 3 is 2.83. The difference of one-hundredth is due to the
rounding of those two probabilities.
The 83 percent probability reflects the drops still falling from the ceiling of the floor below onto the plastic sheets. The notice names no injury. Level 2 still has a share because the same email also has the other reading: the night shift is out of the server room, the early staff on the floor below have gone home, the valve handle is fully over, and the pipe on the engineer’s side of the valve is quiet. I would put this email near the front of the overnight queue, ahead of a notice that has people out of their rooms only and a source described as held. The cabinets in the server room will matter to the reserve and to any sublimit, and that is done with the sums in the claim system. The peril and the choice between building coverage and a liability claim from the tenants are not addressed in this question. This score orders only the queue.
Score: 2.82
Confidence: 0.82 Confidence is 82 percent. It is a separate field from the level probabilities. Each level probability is a bar, and confidence is one number for the pattern of those bars. A high value means the bars are uneven. It is not the height of the tallest bar, nor is it a secondary probability that the level is correct.
The legend echoes the criteria in the order supplied, from 0 upward. Each level has the probability the model assigned to that index.
Probabilities:
- 0: 0,
- 1: 0,
- 2: 0.17,
- 3: 0.83
JSON response from TypeSafe UI:
{
"model": "jev-1.13.0",
"answers": {
"queue": {
"type": "score",
"score": 2.82,
"legend": {
"0": "The event is over. No occupant is out of use, and no injury is mentioned.",
"1": "Damage is at the insured premises only. The insured can keep operating. No injury is mentioned.",
"2": "Another occupant is out of use, or the insured's operations have stopped.
The narrative says the source is under control. No injury is mentioned.",
"3": "Water, fire, or collapse is still in progress, or the narrative names an injury."
},
"confidence": 0.82,
"probabilities": {
"0": 0,
"1": 0,
"2": 0.17,
"3": 0.83
},
"stats": {}
}
},
"usage": {
"input_tokens": 719,
"output_tokens": 17
},
"request_id": "playground_1e306c1396c6c4b404fa7268963c45e9140",
"evaluation_time_ms": 52.74227600602899
}
Noul: Identifying Underwriting Exclusions
Noul is the simplest primitive of the three. You state one proposition. Jev returns one number, noul, the probability that the answer is yes.
A value near 1 represents a strong yes, near 0 a strong no, and near 0.5 that the model is uncertain. There is no confidence field in a Noul response.
Use it when the underwriting question is whether a condition is present, rather than which class should win. A Choice always spends its probability inside the list. A Noul can come back low. This narrative never uses the word “recycling.” The physical operation is still material recovery, which is the class many property guidelines restrict or decline. Underwriting appetite is determined by the carrier; Noul merely determines the factual presence of the condition. The guideline, including the euphemisms that have shown up on slips, lives in the criteria.
Jev Request and Resulting Response
Input
State:
{
"state": "The insured runs a regional consolidation centre for retail packaging. Mixed post-consumer plastic and cardboard arrive from stores, are sorted on site,
baled, and sold on to reprocessors. A small office overlooks the yard."
}
Questions:
{
"material_recovery": {
"type": "noul",
"instructions": "Is this location sorting, baling, or otherwise processing post-consumer waste or secondary raw materials?",
"criteria": {
"true": "The location sorts, bales, shreds, or otherwise processes post-consumer waste, scrap, or secondary raw materials. Wording such as consolidation centre,
logistics hub, or packaging recovery still counts when that processing is described.",
"false": "The location stores or hauls intact finished goods and does not open, sort, or bale a waste stream."
}
}
}
The questions carry the instructions and the criteria, which are true and false. The state is a separate field.
Response

A score of 0.99 is a strong yes: the narrative says the goods are post-consumer, sorted, baled, and sold to reprocessors. It is one draw. A bucket of files tagged 0.99 should be mostly right, and it will contain misses; therefore, 0.99 is an ill-advised threshold for declining a submission without human review unless the carrier has measured that band on its own slips. I would auto-decline only on a higher bar, and only together with a second Noul confirming that the operation is described specifically enough to classify. A thin “logistics center” with no process described should be referred, including when the exclusion number is the lower one. A pure cross-dock of intact cartons is the false case in the criteria, and it should sit near 0.
The same request can carry several of these questions, and a workflow could maintain one Noul per excluded operation and one per named hazard, all on the same passage. PML, capacity, policy limits, and any other numeric thresholds remain in ordinary code. An uncertain probability is referred to an underwriter. Auto-decline, if a carrier uses it at all, requires a high yes-probability on that one exclusion plus a separate Noul confirming that the description is specific enough to classify. Both threshold cutoffs belong to the carrier. One request against one state is the cost-effective call. Output tokens are not billed. Published model-call times are sub-second, and the figure moves with the payload and with where the caller is located. Parsing the document falls outside that figure.
JSON response from TypeSafe UI:
{
"model": "jev-1.13.0",
"answers": {
"material_recovery": {
"type": "noul",
"noul": 0.99,
"stats": {}
}
},
"usage": {
"input_tokens": 424,
"output_tokens": 23
},
"request_id": "playground_1e3a9952ddde54a471fa0e481877b449990",
"evaluation_time_ms": 38.57667299962486
}
Conclusion
The occupancy Choice above is a short book. I have seen occupancy lists with more than 500 rows. A Choice holds at most 255 options, and the published classification guidance puts reliable behavior closer to 240, so a list of that size is a hierarchy of Choice questions. Every company also sets its own PML by line of business and its own capacity by type of risk, and those comparisons are executed in code. An underwriting guideline can run to hundreds of pages. That is why Jev cannot be used on its own. TypeSafe describes Jev as a machine-to-machine model. It fits into workflows that also use a large language model: the language model prepares or drafts the text, and Jev returns the decision on the text it is handed.
FAQ
Who built Jev?
Jev was built by the artificial intelligence startup TypeSafe, founded by Diogo Almeida, Erik Gafni, and Sasha Sheng. The model was released in early access on September 15, 2026, and the current version is 1.13.
What is Jev?
Jev is a System One model. It makes a fast judgment on a state you provide and returns a strictly typed decision for each question you send alongside that state. The state is the text. The questions are separate. Each answer is a Choice, a Score, or a Noul, with the probabilities described below.
What is the state?
The state is the text Jev judges. It can be a string, a JSON object, or an array of text. Images, audio, and video are refused. Questions are a separate field on the request.
What does each primitive return?
Choice is one option from an unordered closed set. The return values are choice, probabilities, and confidence.
Score is a position on an ordered rubric you write, least to most, with 2 to 10 levels, indexed from 0. The return values are score, legend, probabilities, and confidence. The score is the probability-weighted mean of the level indexes and can fall between two levels.
Noul is the answer to “Is this statement true?” The return value is one float from 0 to 1, the probability that the answer is yes. Noul does not return a confidence value.
What Are the Fundamental Differences Between Jev and LLMs in Insurance Workflows?
Large language models read text and produce text, token by token. A workflow then has to parse a label out of that prose, and the prose can be wrong. Jev is a non-generative decision model. It returns one of three shapes: a Choice among options you set, a Score on levels you set, or a Noul, the probability of yes. There is no generated sentence, so there is nothing to parse, and an insurance system or a language model alongside it can read the fields directly. The answer stays inside the schema you declared. The selected option can still be the wrong one, and the same call can move slightly from one run to the next, so a gate should be a band. Several triage questions are scored in one pass. Published model-call times are sub-second and move with the payload and with where the caller is located. Input pricing is $0.042 per million tokens. Output tokens are not billed, which is much cheaper than paying for generated text. Jev reads the question shape and returns the matching answer shape.
What are the top agentic AI use cases for Jev in insurance?
TypeSafe’s use-case map names insurance claims and risk assessment as places to try this, and it includes underwriting workflows in one line about turning narratives into probabilistic indicators. That map represents an exploratory brainstorm. There is no published underwriting-submission product behind it. What follows is my reading of where the three primitives fit.
On a first notice of loss, a Choice can route a narrative to a line of business, and a Score can order the queue on levels you describe. The loss amount stays in code. Straight-through processing is also code. A payout is allowed only when the semantic checks fall within the acceptance band, and the amount, the dates, and the limits pass in the policy system. Jev supplies the semantic checks. The policy system owns the payment.
In underwriting, I would use it for submission appetite triage after the package has been reduced to the passages that matter: one Noul per excluded operation or named hazard, and a Choice over the occupancy list the tariff actually uses, with an explicit “not stated” option. A person reviews the uncertain band. A Score on one dimension can order the queue. The fast lane stays a conjunction in the carrier’s own code.
Urgency and escalation are the same pattern on incoming policyholder text. A Score says how urgent the message is, and the code flags whatever clears the band you set. The fit I see is classification, detection, and prioritization of text, plus probabilities used as features for a risk model the carrier already governs. Pricing, PML, capacity, and authority limits stay outside the model.
Which operational risk controls should be implemented when delegating insurance cases to Jev?
Bands set on your own labeled files. For a Noul, band the yes-probability. A probability near 0.5 reflects uncertainty and should be routed to a human underwriter. Auto-decline, if you use it, requires a high yes-probability on one exclusion plus a separate specificity check, and both threshold cutoffs are determined by the carrier. For Choice and Score, read the winner’s probability and the confidence. Confidence measures how peaked the distribution is. Calibration, as TypeSafe states it, is a group property of the probabilities: a bucket tagged 0.8 should be right about 80 percent of the time across many predictions, and a single case at 0.8 can still be the wrong one. A peaked, low-impact case can go straight through only when the rest of the code gates pass. The middle band goes to a person.
Arithmetic stays in code. Premiums, deductibles, limits, depreciation, PML, capacity, and total insured values are ordinary calculations.
Closed-option guardrails. A categorization schema needs an explicit needs_review or other option, so that a vague file is directed into a review bucket you defined.
Is Jev good in all languages?
No. English is the primary training language, and accuracy on English text is currently best. Other languages are supported as well, though they are not equally accurate. No German benchmark is published. Insurance companies working in English can start an MVP in the language the model was trained for, and anyone else should measure the referral band on their own files first.
How to ensure auditability and reduce the general risk of using Jev in insurance?
Keep an audit trail. Log the request ID, the question names, the instructions, the criteria versions, the probabilities, and the threshold the code applied. Keep the raw state out of the log. Submissions and claim notes contain customer data.
Jev does no chain-of-thought reasoning and returns no built-in rationale. For a Choice or a Score, the model produces the distribution over the options or levels you supplied. For a Noul, it produces the single yes-probability. The record a person can follow is that result and the path the code took.
Document manipulation is a risk. Fabricated or misleading text can shift the probabilities. The model does not treat the state as hostile. Validate documents in an independent layer before they become state.
Sources
- TypeSafe: To run any of the examples in TypeSafe’s playground, you must create an account. Then navigate to the billing tab, add a payment method, and buy at least $5 of credit. This should be sufficient to explore, because Jev is much more cost-effective than any LLM.