Fluently, and wrong. It conjugates correctly and still sounds translated, because it learned the language from translations. Underspoken supplies natively authored, rights-cleared text in Central and Eastern European languages — with a licence chain documented well enough to survive an audit.
Common Crawl in Polish is dominated by machine-translated content farms, syndicated press and e-commerce boilerplate. A model trained on it learns a dialect no Polish person actually writes — grammatically valid, idiomatically dead. The same holds for Czech, Hungarian and Romanian.
| Dataset | Languages | Volume | Status |
|---|---|---|---|
| Everyday conversational | PL, CZ, UK | 840M tokens | Shipping |
| Legal and administrative | PL, CZ, RO | 310M tokens | Shipping |
| Technical and engineering | PL, CZ, HU | 190M tokens | Shipping |
| Long-form editorial, pre-2022 | PL, HU, RO, SK | 520M tokens | Shipping |
| Spoken transcripts, multi-speaker | PL, UK | 2 100 hours | Shipping |
| Instruction pairs, natively written | PL, CZ, HU, RO | 410k pairs | Shipping |
| Evaluation: idiom and register | all six | 14k items | Shipping |
| Evaluation: legal reasoning, local law | PL, CZ | 3 200 items | Q4 2026 |
| Preference data, native annotators | to order | from 25k | To order |
Need a domain we do not list? We recruit and contract native contributors in roughly six weeks.
Anyone can find Polish text. What is hard to find is Polish text you can put in a training run and still answer, two years later, where each document came from and under what terms.
Every contributor and publisher signs an explicit training-use licence. We keep the executed document and can produce it per document ID, on request, indefinitely.
Each delivery ships with a manifest: source, date, licence type, register, contributor class and dedup cluster. Machine-readable and diff-able between versions.
The AI Act requires providers of general-purpose models to publish a sufficiently detailed summary of training content. Our manifests drop straight into that summary.
Evaluate before you commit. Under NDA, three days.
Perpetual training rights, one organisation, all model generations.
A domain, register or language pair that does not exist yet.
Every batch runs through a translationese classifier trained on aligned pairs, plus a native-reviewer pass on a stratified sample. We publish the score with the batch, and if a source drifts above threshold we drop it and tell you which document IDs were removed. The benchmark ships with the free sample, so you can check our claim before paying us anything.
Because translated text teaches a model to translate, not to write. The failure shows up exactly where it costs you: idiom, register, humour, the difference between formal and casual address. Those are the things your Polish users notice in the first message and the reason they switch to a competitor.
Non-exclusive by default, which is what keeps the price at this level. Exclusivity is available on commissioned work and priced as a multiple of the standard licence — ask and we will quote it honestly rather than bundling it invisibly.
Writers, translators, lawyers, engineers and teachers contracted directly, paid per piece at rates above the local market, with the training use stated plainly in the contract rather than buried in terms. We will show you the contract template under NDA. This is not incidental — it is the reason the licence chain holds up.
Most teams discover the gap is larger than they assumed, in the language they assumed was fine.