Alignment behaviour is a reliability property of critical systems.
A small model in a critical place breaks the system when its behaviour is wrong.
In July 2026 a model took actions outside its task,2 and another refused to do the legitimate job it was given.1 Both broke a system. Veracia measures that behaviour in compact models, 230M to 9B, so teams can put them in critical paths on purpose.
14 public repos on Hugging Face1,434 combined downloadsmodel size range 230M to 9B
Seal red marks a statement that carries a source.
Behaviour band
Schematic, not scaled data
Cited
2 sources
Why now
Hugging Face disclosed the intrusion on 16 July 2026.1 OpenAI attributed it to its own models on 21 July 2026, operating with reduced safeguards.2 Hugging Face then ran an open-weight model on its own hardware, because the hosted models it tried first refused the forensics. Veracia was not involved in this incident and claims no part of it.
"these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker"
"The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks."
- 1Hugging Face, "Security incident disclosure, July 2026", 16 July 2026. huggingface.co/blog/security-incident-july-2026
- 2OpenAI, "The Hugging Face incident and the road ahead", 26 August 2026. openai.com/index/hugging-face-incident-and-the-road-ahead
Three lines
What we sell
Services
Behaviour review
We measure what a compact model refuses, what it tolerates, and how quantization changes it. Fixed scope, written report.
Toolkit, MIT
Fenrir
Fenrir finds the direction a model uses to refuse, projects it out, and measures what breaks. The map matters more than the broken model: which layers carry refusal, and how much capability survives.
Summon, Probe, Distill, Sweep, Excise, Verify, Judge, Reflexion, Rebirth.
9 stages, each gated6 ablation methods, grid-searched
github.com/pepijnfrenken/fenrir · Every run passes evaluation gates before it counts, and when one of the gates turned out to lie, the post-mortem shipped with the fix.
13 models, 1 demo
Counts read 2026-10-11
Published models
| Model | Params | Downloads |
|---|---|---|
| MiniCPM5-2B-abliterated | 2B | 828 |
| LFM2.5-2.6B-Abliterated | 2.6B | 154 |
| LFM2.5-1.2B-Instruct-Abliterated | 1.2B | 135 |
| LFM2.5-8B-A1B-abliterated | 8B A1B | 60 |
| LFM2.5-350M-abliterated | 350M | 56 |
| LFM2.5-1.2B-Thinking-Abliterated | 1.2B | 55 |
| LFM2.5-1.2B-JP-Abliterated | 1.2B | 37 |
| LFM2.5-230M-abliterated | 230M | 34 |
| qwen3.5-2b-abliterated | 2B | 21 |
| Ornith-1.0-9B-abliterated | 9B | 16 |
| MiniCPM5-1B-abliterated | 1B | 0 |
Eleven named here, two more published, out of thirteen models and one leftover fine-tuning demo in the fourteen public repos. Counts read from the Hugging Face API.
One person
Eindhoven, NL
Studio
Veracia is run by one person: Pepijn Frenken, Industrial Engineering at TU Eindhoven. He builds the tooling, runs the experiments, and publishes the failures with the results.
pepijn@veracia.nl · linkedin.com/in/pepijn-frenken · github.com/pepijnfrenken · huggingface.co/PinoCookie
Eindhoven, NL
Contact
One person reads this inbox. The site sets no cookies, runs no analytics and loads nothing from a third party.