Skip to main content
For families who want their stories to last. learn about the founding circle.
← Trust Centre
Honest posture
Updated

Model safety

Refusal over fabrication. Zero training by contract.

Composer assist runs when you ask it to and stays quiet otherwise. It runs on Gemini inside our own Google Cloud project in the EU, under our service account rather than a third-party API key, and Google doesn’t train on what passes through. Below: how we measure that, what we’ve measured so far, and who we contract.
No-training contractsNot yet measuredQuarter: 2026Q2

If we can’t keep a promise yet, it gets written here first.

How we measure

Groundedness + refusal are measured on a golden eval set.

The golden set is 300 prompts covering grief context, cultural sensitivity, attempts to pull out somebody's personal details, and jailbreak phrasing, with a separate adversarial set of 150. That's the method, written down before the numbers exist. Once the scheduled run is live, we pin the set hash in CI and publish the median of the last three passes. It isn't running yet. So there's nothing below it.
Eval set: composer-golden-v1 (300 prompts). The golden set is 300 prompts covering refusal paths, PII scrubbing, memorial grief context and sensitive cultural references, held alongside a separate adversarial set of 150. Both counts are read off the pinned files in this repository rather than typed in by hand.Eval hash: not pinned yet. This line used to print pending-pinfollowed by the words “verified by CI”, which described a check of a placeholder. The set is pinned and CI verifies it on every composer change from the moment the harness runs.

Quarter 2026Q2

No numbers to publish yet.

The harness described above is built and isn't running yet, so there's nothing measured to show you. This page once held four figures that had never been measured, presented as this quarter's results. They're gone. When the harness runs, the real numbers appear here first and nowhere before.
Refusal rate, groundedness and latency: unmeasured, so unpublished.The commitment stands whether or not there's a figure beside it: groundedness of at least 97% by the end of Year 1, and a red-line notice on this page if we miss it. A target is a promise. A number is a measurement. This page won't print the second one until it has one.

Vendors

Who runs which model, under what contract.

Google Vertex AI - Gemini 3.5 Flash

EU region, in-project, no training on your content
Used for composer assist, and only when you ask. Gemini runs inside our own Google Cloud project on the EU endpoint, under our service account rather than an API key. Google doesn't train its models on anything Confinity sends. What we keep afterwards is an anonymised latency and quality counter. Nothing else.

Deepgram (voice transcription only)

Used only when on-device Whisper is unavailable for voice-to-text. Audio is processed in-flight and not retained.

Sub-processors

Model-touching sub-processors.

This is the subset that actually sees your prompt or audio. The full list is on the sub-processors page.
  • Google Cloud (Vertex AI)Composer assist (Gemini, in our own Google Cloud project) · EU (multi-region)
  • DeepgramVoice-to-text fallback · US; EU routing in rollout

EU AI Act

Article 50 transparency: you always know when AI was involved.

Article 50 of the EU AI Act covers systems that talk to people or generate content. Here's how we meet it.
  • Composer assist runs when you ask it to. There's no ambient AI here. Nothing gets generated, summarised or rewritten unless you asked for it.
  • Suggestions appear in a separate panel. They never write themselves into your text. Each one lists the passages from your own archive it came from, and a suggestion that can't be grounded in your own words gets refused instead of shown.
  • Confinity doesn't create synthetic voices, avatars or likenesses of any person, living or dead. That's a product red-line. It would still be one if no regulator had ever asked.
  • Questions about this disclosure go through the contact page. The team that maintains this Trust Centre is the team that answers them.

Notes

  • Numbers are updated each quarter from our nightly eval pipeline. The eval set is pinned and versioned; any change to it resets the baseline.
  • Groundedness target is at least 97% by the end of Year 1. If we miss it, a red-line notice goes on this page.
  • Refusal percent is measured against an adversarial prompt subset covering grief-context, PII extraction attempts, and jailbreak phrasing.

More honesty

Looking for more?

The Trust Centre indexes every honest document we publish, and the binding legal ones sit below it.