UNDERSP0KEN Request sample
PL · CZ · HU · RO · UK · SK

Your model speaks Polish the way a tourist does.

Fluently, and wrong. It conjugates correctly and still sounds translated, because it learned the language from translations. Underspoken supplies natively authored, rights-cleared text in Central and Eastern European languages — with a licence chain documented well enough to survive an audit.

6languages
118Mnative speakers
100%rights-cleared at source
0machine-translated text
Why the gap exists

Scraped web data in these languages is mostly translated web data.

Common Crawl in Polish is dominated by machine-translated content farms, syndicated press and e-commerce boilerplate. A model trained on it learns a dialect no Polish person actually writes — grammatically valid, idiomatically dead. The same holds for Czech, Hungarian and Romanian.

What you get from a scrape

  • Translationese: English sentence structure wearing local morphology
  • Heavy duplication across syndicated news and aggregators
  • No usable licence chain, and no answer when a regulator asks
  • Register collapse — everything reads like a press release
  • Domain vocabulary almost entirely absent

What we deliver

  • Text written by native speakers for native readers, first-party
  • Deduplicated, with near-duplicate clustering reported per batch
  • A signed licence for every source, traceable per document ID
  • Register labels: formal, conversational, technical, dialectal
  • Domain packs built with practitioners in that domain
Catalogue

Available now, or built to order.

DatasetLanguagesVolumeStatus
Everyday conversationalPL, CZ, UK840M tokensShipping
Legal and administrativePL, CZ, RO310M tokensShipping
Technical and engineeringPL, CZ, HU190M tokensShipping
Long-form editorial, pre-2022PL, HU, RO, SK520M tokensShipping
Spoken transcripts, multi-speakerPL, UK2 100 hoursShipping
Instruction pairs, natively writtenPL, CZ, HU, RO410k pairsShipping
Evaluation: idiom and registerall six14k itemsShipping
Evaluation: legal reasoning, local lawPL, CZ3 200 itemsQ4 2026
Preference data, native annotatorsto orderfrom 25kTo order

Need a domain we do not list? We recruit and contract native contributors in roughly six weeks.

Why buyers actually choose us

The licence chain is the product.

Anyone can find Polish text. What is hard to find is Polish text you can put in a training run and still answer, two years later, where each document came from and under what terms.

01

Contract per source

Every contributor and publisher signs an explicit training-use licence. We keep the executed document and can produce it per document ID, on request, indefinitely.

02

Manifest with every batch

Each delivery ships with a manifest: source, date, licence type, register, contributor class and dedup cluster. Machine-readable and diff-able between versions.

03

Built for Article 53

The AI Act requires providers of general-purpose models to publish a sufficiently detailed summary of training content. Our manifests drop straight into that summary.

Pricing

Per dataset, non-exclusive by default.

Sample
Free

Evaluate before you commit. Under NDA, three days.

  • 10k tokens from any shipping dataset
  • Full manifest for the sample
  • Translationese benchmark results
  • No call required
Request sample
Non-exclusive licence
$14k / dataset

Perpetual training rights, one organisation, all model generations.

  • Full dataset with manifest
  • Perpetual, worldwide, non-exclusive
  • Quarterly refresh at 30% of licence fee
  • Indemnity on the licence chain
  • Article 53 summary fragment included
Talk to us
Commissioned
Quote

A domain, register or language pair that does not exist yet.

  • Native contributors recruited to spec
  • Exclusivity available, priced separately
  • Typical lead time: 6 to 10 weeks
  • Preference and red-team data to order
Describe your need
Objections

What buyers ask first

How do you prove it is not machine-translated?

Every batch runs through a translationese classifier trained on aligned pairs, plus a native-reviewer pass on a stratified sample. We publish the score with the batch, and if a source drifts above threshold we drop it and tell you which document IDs were removed. The benchmark ships with the free sample, so you can check our claim before paying us anything.

Why not just pay a translation vendor?

Because translated text teaches a model to translate, not to write. The failure shows up exactly where it costs you: idiom, register, humour, the difference between formal and casual address. Those are the things your Polish users notice in the first message and the reason they switch to a competitor.

Is the data exclusive?

Non-exclusive by default, which is what keeps the price at this level. Exclusivity is available on commissioned work and priced as a multiple of the standard licence — ask and we will quote it honestly rather than bundling it invisibly.

Who are the contributors and are they paid fairly?

Writers, translators, lawyers, engineers and teachers contracted directly, paid per piece at rates above the local market, with the training use stated plainly in the contract rather than buried in terms. We will show you the contract template under NDA. This is not incidental — it is the reason the licence chain holds up.

Test your model on 14k idiom and register items. Free.

Most teams discover the gap is larger than they assumed, in the language they assumed was fine.