Release
llama.cpp server adds support for decision models through a /v1/systemone endpoint
On October 2, 2026 the ggml team announced on the Hugging Face blog that the llama.cpp server now supports decision models through a new /v1/systemone endpoint. Decision models score the options a caller supplies instead of generating text. The pull request that adds the endpoint (#29818) was merged into the llama.cpp repository the same day; it describes these models as wrappers around existing embedding-model architectures and adds a decision head and server handling for them. The announcement lists six supported models (Julia-1, Laya, Kev-4B, lev, OpenJev, and Clef), while the pull request itself covers five of them and lists Clef support as a planned follow-up.12