Local models
Quantised open-weight models on consumer hardware. What fits in 12 GB, what it costs in quality, and where the ceiling really is.
Gribbite is an independent technology lab. I run models on my own hardware, wire up agents, automate the tedious parts of cloud operations, and then test how all of it behaves when something hostile shows up. The work gets built, broken, and written down here.
Six threads, all of which end up touching each other. Nothing here is theoretical — each one exists because I needed to know how something actually behaves.
Quantised open-weight models on consumer hardware. What fits in 12 GB, what it costs in quality, and where the ceiling really is.
Small, scoped agents that do one job well: tool use, retrieval, and the plumbing that keeps them from wandering off.
Prompt injection, tool-call abuse, data exfiltration through model context. Treating an agent as a workload that needs a threat model.
Writing, tuning, and breaking detections in KQL. Testing whether a rule survives contact with noisy production telemetry.
Entra ID governance, access packages, lifecycle automation, and the failure modes that only show up at scale.
PowerShell and Graph automation for the work that shouldn't need a human twice. Scripts that other people can read.
Each project page carries the setup, what was measured, and what turned out to be wrong.
active
A repeatable rig for loading, quantising, and comparing open-weight models on a single 12 GB card — and finding the point where a smaller model stops being worth it.
ongoing
Giving a tool-using agent real permissions in a throwaway tenant, then trying to talk it into doing something it shouldn't. Notes on what actually held.
active
A small harness for writing KQL detections against simulated activity, measuring false positives before a rule ever reaches production.
The reference material I keep going back to, filtered down to things worth your time.
Inference runtimes, quantisation formats, and where to find weights that are actually licensed for what you want to do.
Threat models, red-team frameworks, and risk guidance that predates the current wave of vendor marketing.
Query languages, detection rule libraries, and the identity documentation worth reading end to end.
If you're building with local models, hardening agents, or writing detections, I'm happy to compare notes.