Veracia
VeraciaEindhoven, NLRev H2026-10-11

Alignment behaviour is a reliability property of critical systems.

A small model in a critical place breaks the system when its behaviour is wrong.

In July 2026 a model took actions outside its task,2 and another refused to do the legitimate job it was given.1 Both broke a system. Veracia measures that behaviour in compact models, 230M to 9B, so teams can put them in critical paths on purpose.

Sourced marks
Specimen size

14 public repos on Hugging Face1,434 combined downloadsmodel size range 230M to 9B

Seal red marks a statement that carries a source.

Figure 1
Behaviour band
Schematic, not scaled data
Figure 1 · Behaviour bandSchematic
OVER-REFUSAL WORKING BAND MISALIGNMENT Lockout Acts outside its task the model will not do the legitimate job stays inside the task the model takes actions beyond the brief [1] INCIDENT RESPONSE LOCKED OUT [2] AGENT ACTED OUTSIDE ITS TASK
Schematic, not scaled data. The band sets the two failure directions a behaviour review has to bound, and both marked ends carry a source quoted in its own words.
Why now
Cited
2 sources

Why now

Hugging Face disclosed the intrusion on 16 July 2026.1 OpenAI attributed it to its own models on 21 July 2026, operating with reduced safeguards.2 Hugging Face then ran an open-weight model on its own hardware, because the hosted models it tried first refused the forensics. Veracia was not involved in this incident and claims no part of it.

Left end · over-refusal
"these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker"
Hugging Face, on its own incident response.1
Right end · misalignment
"The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks."
OpenAI, describing its own models.2
  1. 1Hugging Face, "Security incident disclosure, July 2026", 16 July 2026. huggingface.co/blog/security-incident-july-2026
  2. 2OpenAI, "The Hugging Face incident and the road ahead", 26 August 2026. openai.com/index/hugging-face-incident-and-the-road-ahead
Services
Three lines
What we sell

Services

Behaviour review

We measure what a compact model refuses, what it tolerates, and how quantization changes it. Fixed scope, written report.

Fenrir
Toolkit, MIT

Fenrir

Fenrir finds the direction a model uses to refuse, projects it out, and measures what breaks. The map matters more than the broken model: which layers carry refusal, and how much capability survives.

Summon, Probe, Distill, Sweep, Excise, Verify, Judge, Reflexion, Rebirth.

9 stages, each gated6 ablation methods, grid-searched

github.com/pepijnfrenken/fenrir · Every run passes evaluation gates before it counts, and when one of the gates turned out to lie, the post-mortem shipped with the fix.

Index
13 models, 1 demo
Counts read 2026-10-11

Published models

ModelParamsDownloads
MiniCPM5-2B-abliterated2B828
LFM2.5-2.6B-Abliterated2.6B154
LFM2.5-1.2B-Instruct-Abliterated1.2B135
LFM2.5-8B-A1B-abliterated8B A1B60
LFM2.5-350M-abliterated350M56
LFM2.5-1.2B-Thinking-Abliterated1.2B55
LFM2.5-1.2B-JP-Abliterated1.2B37
LFM2.5-230M-abliterated230M34
qwen3.5-2b-abliterated2B21
Ornith-1.0-9B-abliterated9B16
MiniCPM5-1B-abliterated1B0

Eleven named here, two more published, out of thirteen models and one leftover fine-tuning demo in the fourteen public repos. Counts read from the Hugging Face API.

Studio
One person
Eindhoven, NL

Studio

Veracia is run by one person: Pepijn Frenken, Industrial Engineering at TU Eindhoven. He builds the tooling, runs the experiments, and publishes the failures with the results.

pepijn@veracia.nl · linkedin.com/in/pepijn-frenken · github.com/pepijnfrenken · huggingface.co/PinoCookie

Mail
Eindhoven, NL

Contact

pepijn@veracia.nl

One person reads this inbox. The site sets no cookies, runs no analytics and loads nothing from a third party.