Data & Transparency
We say "we publish our accuracy measurements"; this page is what that promise is worth. The model's report card measured on a real feed, how the measurement was done, the criteria the labels follow, and where the card should not be trusted.
Report card
The numbers below were measured on a real daily feed of 1,110 headlines that never appeared in the model's training, using the thresholds the app actually ships with. The values are given per language: the model behaves differently in each, and since 92% of the measurement set is Turkish, a combined table would present the Turkish numbers as the overall result. You can recompute the same numbers yourself from the downloadable data.
Turkishn = 1,024
| Task | Positives | P | R | F1 | TP / FP / FN |
|---|---|---|---|---|---|
| Critical | 52 | 0.90 | 0.52 | 0.659 | 27 / 3 / 25 |
| Low quality | 150 | 0.99 | 0.61 | 0.752 | 91 / 1 / 59 |
| Sentiment · negative | 285 | 0.76 | 0.80 | 0.783 | 229 / 71 / 56 |
| Sentiment · neutral | 598 | 0.83 | 0.86 | 0.845 | 516 / 107 / 82 |
| Sentiment · positive | 141 | 0.67 | 0.48 | 0.562 | 68 / 33 / 73 |
Sentiment macro-F1 0.730 · accuracy 0.794
Englishn = 86
| Task | Positives | P | R | F1 | TP / FP / FN |
|---|---|---|---|---|---|
| Critical | 15 | 0.87 | 0.87 | 0.867 | 13 / 2 / 2 |
| Low quality | 7 | 0.75 | 0.43 | 0.545 | 3 / 1 / 4 |
| Sentiment · negative | 35 | 0.75 | 0.69 | 0.716 | 24 / 8 / 11 |
| Sentiment · neutral | 38 | 0.68 | 0.84 | 0.753 | 32 / 15 / 6 |
| Sentiment · positive | 13 | 0.86 | 0.46 | 0.600 | 6 / 1 / 7 |
Sentiment macro-F1 0.690 · accuracy 0.721
Only 7.7% of the measurement set is English. The low-quality row rests on seven positive examples and the positive-sentiment row on thirteen; a single item moving shifts those ratios substantially. Do not read this block as a reliable estimate.
P (precision) is how much of what we labelled was correct, R (recall) is how much of what should have been labelled we caught, F1 is the balance between the two. Sentiment is not compressed into a single average: each class gets its own row, measured as "this class or not". That keeps the weak spot visible — recall on positive sentiment is markedly low, meaning we call nearly half of the good news neutral. TP is correctly caught, FP wrongly labelled, FN missed.
Why is recall so low?
Because we chose it to be. MagPunk is tuned on the principle that missing a label is better than pressing a wrong one: the decision thresholds are not calibrated to maximise F1, but to keep precision above a floor and maximise recall subject to that constraint.
The floors are: critical TR 0.80 · EN 0.90, low quality TR 0.90 · EN 0.90. So the low R in the table is not a failure, it is the price we knowingly pay. Wrongly showing an item as "critical" does more damage than missing a critical one.
Could this table flatter the model?
It could — which is why we do not measure there. The training pool is enriched by mining: minority classes appear far more often than they naturally would. In a real feed the share of critical items is about 6%; in the English side of the training pool it rises to 16.7%, roughly 2.8 times the real rate. The higher the base rate, the easier precision becomes.
That is why the report card above is measured not on the enriched pool but on a 1,110-headline slice taken from a real feed that never entered training. The numbers on the enriched pool look prettier; we chose not to put those in the window.
How the measurement was done
The measurement set was separated from a real daily feed, contains 1,110 headlines, and does not intersect the training pools — overlapping rows were removed from the honest benchmark.
The decision thresholds were calibrated not on this set but on a further separate validation slice, so the thresholds were not tuned to the data they are scored on.
The model reads the title only and never looks at the body. This is a deliberate privacy and speed choice, but it is also an information ceiling: a misleading title misleads the model too.
Where the training data came from
The headlines the model was trained on were compiled from news-headline datasets published publicly for research. Nothing from the app's own feeds or from users' reading data was used — which is also technically impossible, because that data never leaves your device.
This has a consequence for reading the card: the training text comes from research corpora while the measurement comes from a real daily feed. The two distributions are not identical; the model was trained in a slightly different world of text from the one the app sees every day.
Most of these datasets grant no redistribution permission. That is why we do not publish the training pool as a download.
The limits of this report card
The ground truth currently comes from a single model. The labels used as "the right answer" were produced by a large language model; how much two human annotators would agree with each other has not been measured yet. So for now this card answers "how well does the model match a consistent teacher", not "how well does the model match a human".
That does not invalidate the numbers, but it bounds them. Work on label consistency is ongoing: the plan is to have a few hundred headlines labelled by two independent people and to publish the agreement between them (Cohen's κ). When that measurement is complete, the result will be added to this page — whatever number it turns out to be.
A second limit: the card covers only the critical, low-quality and sentiment labels. Rule-based components such as the score, the similarity threshold and first-publisher detection are not measured in this table.
What criteria the labels follow
Labels are not given by intuition but by a written decision procedure. The procedure is sequential: once a step decides, the next is not reached. The reason is the typical failure seen in earlier rounds — fixating on a single word in the title and missing the context.
First filter: is this news, or a query?
If the title is a question, a service query or a list — "watch live", "how much", "how did they die", "how to", lists starting with a number — it is marked low quality and the procedure stops there. This step exists precisely for cases like "Was there an earthquake in Yalova?", which is not an earthquake report but a query; it does not get the critical label.
Critical: did the event actually happen, and is it societal?
A natural disaster that has occurred, or a mass event with loss of life, is critical; mass death is always critical.
Not critical: the outcome of a security operation (detentions, seizures), individual events (a single traffic accident or murder), commemorations and anniversaries (the event is not happening now), and metaphorical usage.
Sentiment: a single question is asked
The question is: what did this event change compared to yesterday? Sentiment looks at the outcome of the event, not at whether individual words are positive or negative. If new harm arises it is negative; if existing harm ends or a grievance is resolved it is positive; if a process is running with no clear outcome it is neutral. That is why "a theft was committed" is negative and "the thief was caught" is positive.
A crime being solved counts as positive only under two strict conditions: the threat is a universal crime (murder, drugs, abduction, child abuse) and the title states a completed resolution ("busted", "rescued", "caught"). A request for arrest, referral to court or the opening of an investigation is a process; it is neutral.
Political neutrality — the strictest rule
All operation and detention news with a political, geopolitical or ideological context, universal crimes aside, is strictly labelled neutral. The reason: calling the routine operation of the state "good" or "bad" would teach the model a political judgement. The only exception is explicitly editorialising language in the title itself.
In sport, routine is separated from the peak
The result of a league match — whoever won — is neutral; it is routine zero-sum competition. A championship, a medal, a cup and their celebration are positive.
A few more distinctions
Irony is resolved: "great, another price hike" is negative. A record is not good news; "record heat" and "record hike" are negative. Surface words like "detention", "court" and "operation" carry no sentiment on their own. Wherever there is doubt the answer is neutral, and that row is marked low confidence.
No title is labelled in bulk by pattern matching; each one is read individually.
A label we removed
Earlier versions had a label called "suspicious", meant to catch curiosity gaps and clickbait. Measured on a real feed its precision came out at 0.12: to catch three pieces of clickbait it was hiding twenty-nine innocent items. It was removed entirely as a result of that measurement — we deleted a label because we measured it.
Downloads
You do not have to take any number on this card on trust. The first file below recomputes every table above.
- Report card data (CSV)35 KB · 1,110 rows
Ground truth and the model's prediction for every item in the measurement set. All the tables above can be recomputed from these two columns alone. It contains no headline text: since P, R and F1 are pure functions of the confusion matrix, verification does not need the headlines — and we do not end up republishing publishers' content.
- Classifier model (JSON)6.1 MB
Byte for byte the file the app loads at startup — not a description of it, the thing itself. It holds the vocabulary, the coefficients, the language-specific decision thresholds and the neutral margins. It contains no headline text; the longest unit it stores is a two-word fragment.
For each task and language, the terms that most influence the decision, with their weights. The fastest way to read what the model is looking at.
- Rule dictionary (JSON)52 KB · 2,718 entries
The hand-written half of the engine. Critical keywords and the list that blocks them (drill, commemoration, anniversary, "did it happen"), sentiment words weighted from −3 to +3, negation and adversative conjunctions, the stop words removed before the similarity calculation, and the Turkish suffix list. Same content as the file the app loads, only broken into lines so it can be read. Before publishing, 139 entries left over from a feature removed from the code — affecting no label — were deleted.
Not published yet
Transparency also means not presenting as published what is not. The following are absent from this page; we give the reasons.
- The training pool itself — most of the datasets it was compiled from grant no redistribution permission.
- Agreement between two human annotators (Cohen's κ) — work is ongoing, and the result will be added here whatever number it turns out to be.
- Any source-level table — we do not publish this as a matter of principle: a label is given to a single title, never to a source.