Skip to main content

ChindaMT: Thai-English translation that follows your rules, published at AACL-IJCNLP 2026

· 6 min read
Kobkrit Viriyayudhakorn
CEO @ iApp Technology

ChindaMT, instruction-following Thai-English translation, published at AACL-IJCNLP 2026

Today we are releasing ChindaMT, a family of Thai-English translation models that follow the rules you give them: the term to use, the tone to write in, the length and the format. The research behind it is published at AACL-IJCNLP 2026 (Main Conference). The models, the training data, the evaluation suites and the code are open under Apache 2.0, and ChindaMT-4B is available now as a hosted API.

Try it: ChindaMT Translator · API: documentation and live demo · Paper: arxiv.org/abs/2609.34770 · Weights: huggingface.co/iapp/ChindaMT-4B

A translation can be correct and still wrong for the job​

Take one sentence: "The meeting has been moved to Friday because the boss is busy."

A standard translation gives การประชุมถูกเลื่อนไปเป็นวันศุกร์เนื่องจากเจ้านายยุ่ง. Every word is right, and nobody would type it into a group chat. Ask ChindaMT to "use a casual, friendly tone" and it writes ประชุมย้ายไปวันศุกร์แล้วนะ เพราะเจ้านายยุ่งมาก.

Translation rarely has one answer. Subtitles need the character's voice, papers need a formal register, a brand needs its own style, a contract needs the same term every time. The difficulty is that the harder a model is pushed to follow such rules, the more its translation quality tends to fall. Translation-specialised models translate well but ignore instructions; general instruction-following models obey but translate less well. No open Thai-English model closed that gap.

The answer was in the data​

We built the training data with a method we call Reference-Grounded Data Curation (RGDC). Instead of asking a language model to invent rules and hoping they can be satisfied, RGDC extracts every rule from a good reference translation that already satisfies it. If the reference uses "THB" for บาท, keeps one sentence and stays formal, those become the rules for that example, and they are achievable by construction because a real translation already meets them.

The pipeline starts from 17.85 million English-Thai sentence pairs from ten public corpora. The first phase keeps the 1.75 million pairs that are hardest but still learnable. The second extracts rules from each reference in five categories (content, numbers, style, format, language), regenerates translations under those rules, and keeps only the candidates that pass every rule. The result is Grounded, a dataset of 1.97 million records with 3.7 rules each on average.

Results​

We fine-tuned three sizes, 4B, 2B and 0.8B, and compared each with open translation models of the same size, in both directions, with and without rules. The figure below is how often a judge model preferred ChindaMT's translation when rules were given (50% is a tie):

ChindaMTCompared withPreferred
4BTyphoon-Translate-1.5-4B68.4%
4BTranslateGemma-4B87.2%
4BMiLMMT-46-4B89.5%
2BHY-MT-1.5-1.8B78.6%
2BGemmaX2-28-2B89.9%
0.8BHY-MT-1.5-1.8B67.8%
0.8BGemmaX2-28-2B86.5%

Following rules did not cost translation quality. On the public FLORES-200 benchmark, ChindaMT-4B scored 90.7 on the mean of three quality metrics, the highest among the 4B to 9B models tested. Three native Thai speakers, rating anonymised pairs without knowing which model produced which, preferred ChindaMT-4B in 69% of translations with rules and 67% overall, more than two to one. The same recipe also improved a second model generation at fixed settings, which indicates the gain comes from the data rather than from one particular base model.

Use it​

On the web. The ChindaMT Translator lets you translate with your own rules and compare the result with a translation without them.

As an API. ChindaMT-4B is live at POST https://api.iapp.co.th/v3/store/text/mt/translate, with a batch endpoint for UI strings and table cells. It takes up to 100,000 characters per request and returns a sentence in about 0.4 seconds and five pages in 1.7 seconds. The price is 1 IC per 400 characters of source text; rules are free and failed requests are not charged. Everything is in the API documentation.

curl -X POST https://api.iapp.co.th/v3/store/text/mt/translate \
-H "apikey: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"text": "The meeting has been moved to Friday because the boss is busy.", "source_lang": "en", "target_lang": "th", "rules": ["Use a casual, friendly tone"]}'

On your own hardware. All three sizes are open under Apache 2.0. The 0.8B model is small enough to run on a laptop.

ReleaseLink
Models (Apache 2.0)ChindaMT-4B, ChindaMT-2B, ChindaMT-0.8B
Training data, 1.97M recordsChindaMT-Grounded
Evaluation suitesChindaMT-CoreEval, ChindaMT-BroadEval
Codegithub.com/iapp-technology/ChindaMT-RGDC
Paperarxiv.org/abs/2609.34770

Organisations that need their own terminology built in, or an on-premise deployment for data security, can contact sale@iapp.co.th.

Limits​

ChindaMT covers Thai and English only. It follows the kinds of rules a good reference translation can demonstrate: terminology, tone, length and output format. A large mandatory glossary and strict placeholder tokens are outside what it was trained for. Like any machine translation it can produce fluent but incorrect text, so translations for legal, medical and other high-stakes use should be reviewed by a person; the API marks likely problems in its warnings field.

Acknowledgements​

ChindaMT is the work of the iApp AI Research team, with Prof. Dr. Thanaruk Theeramunkong. We thank SiamAI for computational resources, OpenThai Lab for its support, and the three native Thai speakers who volunteered as raters.

If you could give a translator one rule, what would it be? Try it in the translator and tell us on Discord.