Anonymization

Anonymized before it changes hands.

Data is put through an anonymization process to remove identifiers, and you review the results before any transfer happens. This page explains how that works and what we ask you to keep out of scope.

Identifiers are removed

The anonymization process is designed to detect and remove direct identifiers — names, email addresses, phone numbers, postal addresses, government IDs and employee numbers — across body text, headers, attachments and transcripts before delivery.

Credentials are filtered out

Detectors look for API keys, tokens, passwords, card numbers and bank details so they can be stripped rather than carried forward. No automated process is perfect, which is why you review the output.

You review before anything moves

You receive category-level statistics and sampled before/after excerpts. Nothing transfers until your team has looked at it and signed off in writing.

You control the transfer

Transfer happens through read-only connectors or a push-only endpoint, using standard encryption in transit and at rest. You decide which sources are in scope and when the window opens.

Use is defined in the agreement

Permitted use, restrictions on transfer to third parties, retention and deletion terms are all written into the licensing agreement your counsel reviews before signature.

You keep your originals

You are licensing a copy. Your systems, retention policies, access controls and legal holds are untouched by this process.

Abstract shield dissolving into redacted text blocks

What it's used for

Training and evaluating leading AI models.

Once anonymized, your corpus is used to train and evaluate leading AI models. Real operational business language — threads, transcripts, tickets, documents — is what these systems are short on, and that is why it is worth paying for. Use is limited to that purpose in the licensing agreement: your corpus is not resold, syndicated or published.

Out of scope

What we ask you to keep out of scope.

These categories are not something we want. If one appears in a scoped source, that source is excluded or the affected records are dropped.

  • Patient health records or anything covered by HIPAA
  • Payment card data and bank credentials
  • Government-issued identification numbers
  • Customer PII in identifiable form
  • Employee HR files, compensation and performance records
  • Anything under active legal hold or NDA restriction

Documentation

What your legal and security teams can review.

Data processing schedule

Categories in scope, retention terms and deletion commitments, written into the agreement.

Anonymization methodology

How identifiers are detected and removed, how samples are reviewed, and how issues are escalated.

Transfer details

How data moves, what encryption is used in transit and at rest, and who on your side authorizes it.

Nothing on this page is a warranty or a guarantee of any particular security outcome. Terms governing anonymization, use, retention and liability are set out in the licensing agreement, which your counsel should review before signing.

Bring your security team. We expect the questions.