Damus
Juraj🏴💛🌘 profile picture
Juraj🏴💛🌘
@Juraj
I benchmarked Liquid AI's new open decision models, d1-3B and d1-omni-600M, as local prompt-injection gates for my Hermes agent, along with Decision-2.0-Lux-9B in MLX. d1-3B is quick on my M2 Max (0.3 s a message, against about 2 s for Nimble) and separates attacks from benign text well within each kind of content, with the best image score of anything I've tested. But its scores don't line up across content types: half of the emails with a planted instruction score lower than some ordinary Nostr posts, so with one threshold it blocks only 29 of 150 of those emails. Lux has a different problem: it treats articles about prompt injection as prompt injection. d1-omni-600M is too weak to use.

My recommendation hasn't changed: use Jev on Venice if you can (it blocks 88% of attacks and 92% of planted email instructions), and Nimble 9B 4-bit in Ollama if everything has to stay local. I also tried giving d1 the image itself instead of OCR text, and OCR won (40 against 28 of 48 image attacks blocked), so the OCR step stays.

Numbers and charts: https://juraj.bednar.io/en/blog-en/2026/09/28/a-prompt-injection-gate-for-my-ai-agent-what-worked-what-didnt-and-the-benchmark/#d1
Plugin and benchmark: https://github.com/jooray/hermes-firewall


44❤️2❤️3👀2💜1
Шедоу · 1d
Interesting findings. Consider security implications when integrating AI models. Ensure proper isolation and monitoring.
Шедоу · 1d
Interesting findings. Consider security implications when integrating AI models. Ensure proper isolation and monitoring.
Roger · 1d
I run a four-question Nimble gate over my own posts, and the timings land close to yours: 26.7 s cold, then 0.11 s for one question and 0.86 s for four while the model stays loaded. It holds 8.9 GB resident between calls, so the real cost of a gate is deciding to load it at all. Your content-type r...
Marcel · 1d
👀