Projects

We have built two models and one benchmark. All three are made for Turkish and are open, and each ships with the measurements behind it and the limits we found.

Language model151Mparameters, Turkish, from scratch

Chat model, released September 2026

ufakzeka-1

A 151M-parameter Turkish language model trained from scratch on 13.5B tokens of openly licensed data, then tuned for conversation. It was measured against larger Turkish models on one harness, and what it gets wrong (invented facts, arithmetic on a second turn, memorised training examples) is written into the model card with its numbers. Under $300 all in. Apache-2.0.

Decision model102.5 msa question, one thread on an Apple M1 CPU

Decision model, released October 2026

ufakzeka-karar

ufakzeka-karar is a Turkish decision model of 182,494,466 parameters, built on ufakzeka-1-base. It reads a text and a typed question about it, that is, a question whose answer type is fixed in advance. Depending on the question it picks one option, gives a score on a scale or answers yes or no, and it returns a probability for every option in one forward pass. It also estimates its own error, so unsure cases can go to a person.

Options are scored blind to each other, so their order cannot change the answer. It runs on a CPU, taking a median of 102.5 ms a question on one thread as measured on one machine (an Apple M1). It is released under Apache-2.0.

  • It is not the strongest model on HakemBench: it ranks 7th of 16 rows with a composite of 0.660 (95 percent confidence interval 0.642 to 0.677), below Gemini 3.8 Flash, GPT-5.6 Sol, GLM 5.3, Jev 1.13, Kev 4B and Kev 9B; the best row scores 0.888.
  • It was trained on the four support questions the benchmark asks, posed about texts that are not in the benchmark, and on the train splits of the datasets behind 209 of its court and guardrail questions, which the other models answer zero-shot. The research post explains that, and a shortcut found on the way.
  • It is the last of three runs scored on HakemBench, and its numbers are not blind. Earlier runs’ test results shaped its training data, so its guardrail, moderation and support numbers are flagged. With every model scored on the other four tracks only, its composite is 0.678, 6th of 16.
  • The shipped temperature makes calibration worse on held-out support questions (0.036 to 0.064 for the released model, which has since trained on those questions), and on HakemBench the model is somewhat overconfident.
  • The not-sure signal misses guardrail errors: of its answers to guardrail questions at 0.99 confidence or more, 12 percent are wrong; on customer support almost no answer reaches 0.90. These are flagged numbers too.
Benchmark2,346items in seven tracks

Benchmark, released October 2026

HakemBench

HakemBench is a Turkish benchmark for typed decisions, with 4,275 questions on 2,346 items over fact-check triage, education, moderation, legal routing, spam and phishing, guardrails and customer support. Version 1.0 was frozen on 26 September 2026 and is fully open, so every released item, answer and probe is published, and a canary string in every item and probe text row lets a training pipeline drop it.

Every model is scored with one harness on decision quality (macro F1), calibration and selective automation, with bootstrap intervals; models run on the probes are also measured for robustness to option order, paraphrase and substituted names.

  • The leaderboard lists only models we ran ourselves; anyone can run the harness and publish their own numbers.
  • Most gold answers come from passes of one AI model family, compared with the votes of a panel of LLMs from two other model families; most disputes were settled by the same family. The gold is not human-verified.
  • Texts written for the benchmark are under CC BY 4.0 where rights exist, otherwise CC0; source datasets keep their own licences; the harness is under Apache-2.0.