Short answer: jev-1.13 does not generate text. You POST a piece of text plus a set of named questions to https://aihubmix.com/v1/systemone, and you get back typed answers keyed by your question names — a category, a score, or a probability. Nothing to parse. Below is a working request, the full response format, and the one design mistake that will quietly cost you accuracy.
What Jev is for
Use it when you need a judgment, not a paragraph: routing a support ticket, scoring severity, flagging urgency, gating content. The model page states it plainly — "No text generation, nothing to parse."
Use a normal chat model instead when you need explanations, summaries, or any free-form output. jev cannot produce those.
Before you start
- [ ] Python 3 installed (the standard library is enough)
- [ ] An AiHubMix account and API key
- [ ] Key exported as
AIHUBMIX_API_KEY, never hardcoded in source
You do not need an OpenAI SDK, LangChain, or the requests package.
Step 1 — Know the endpoint
POST https://aihubmix.com/v1/systemone
Authorization: Bearer YOUR_KEY
Content-Type: application/json
This is a model-specific route, not /v1/chat/completions. OpenAI-compatible clients cannot call it. Use plain HTTP.
Step 2 — Choose your question types
There are three, and you can mix them freely in one request:
| Type | Use it for | You supply |
|---|---|---|
choice | Picking one category | criteria as a dict: label → definition |
score | Rating on an ordered scale | criteria as a list, lowest first |
noul | A yes/no judgment | nothing beyond instructions |
Step 3 — Build the request
state is the text to be judged. questions keys are names you invent; the response uses the same names.
payload = {
"model": "jev-1.13",
"state": "Hi, I have been trying to connect my Stripe account for 3 days "
"and it keeps failing. I am losing sales. Please help ASAP.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions",
},
},
"frustration": {
"type": "score",
"instructions": "How frustrated the customer appears",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language",
],
},
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency or time-sensitivity",
},
},
}
Step 4 — Send it with the standard library
import json, os, urllib.error, urllib.request
def ask(payload):
req = urllib.request.Request(
"https://aihubmix.com/v1/systemone",
data=json.dumps(payload).encode(),
headers={"Authorization": "Bearer " + os.environ["AIHUBMIX_API_KEY"],
"Content-Type": "application/json"},
method="POST",
)
try:
with urllib.request.urlopen(req, timeout=60) as r:
return json.loads(r.read())
except urllib.error.HTTPError as e:
raise SystemExit(f"HTTP {e.code}: {e.read().decode(errors='replace')[:500]}")
Keep the request construction inside a function. If you build and send at module level, importing the file re-fires the call and spends tokens.
Step 5 — Read the answers
Each answer stores its value under a key matching its type, so one accessor covers all three:
data = ask(payload)
for name, ans in data["answers"].items():
kind = ans["type"]
print(name, kind, ans[kind], ans.get("confidence"))
print("usage:", data.get("usage"))
Verified output:
department choice billing 0.51
frustration score 1 1
is_urgent noul 1 None
usage: {'input_tokens': 424, 'output_tokens': 73}
Step 6 — Understand the full envelope
The print loop hides useful fields. The raw JSON for a score answer:
{
"type": "score",
"score": 1,
"legend": {"0": "Calm, just stating facts", "1": "Frustrated but civil", "2": "Very angry, strong language"},
"probabilities": {"0": 0, "1": 1, "2": 0},
"confidence": 1
}
legendmaps the index back to your own wording — the integer is self-describing, so you need no lookup table in your code.probabilitiesgives the full distribution, useful for detecting a near-tie.noulhas noconfidencefield. Its probability is the signal.
The top level also returns usage, id, provider (TypeSafe), and the resolved backend version typesafe/jev-1.13-20260917.
The mistake that costs you accuracy
In the request above, department returned billing with a confidence of 0.51. Repeating the identical request returned the same label at 0.38 — the answer was stable, the stated certainty was not.
The cause is in the criteria, not the model. billing is "Payment or subscription issues" and technical is "Bugs or integration problems." A Stripe integration that keeps failing genuinely matches both. The model was reporting an ambiguity that was written into the schema.
Two rules before production:
- Write
choicecriteria that are mutually exclusive. If a human would hesitate between two labels, the model will too. - Set a confidence threshold and route low-confidence results to a human queue instead of accepting them as decisions. A model that admits uncertainty is more valuable than one that hides it.
Pre-launch checklist
- [ ] Key loaded from environment or a
600-mode file your VCS ignores - [ ]
choicecriteria reviewed for overlap - [ ] Confidence threshold defined, with a fallback path below it
- [ ]
noulanswers handled separately — they carry noconfidence - [ ] Unused questions removed (3 questions cost 424 input tokens; 2 cost 355)
- [ ] Non-2xx responses logged with the raw body
FAQ
Can I use the OpenAI SDK? No. /v1/systemone is not an OpenAI-compatible route.
Can jev return a sentence or summary? No. It answers only the typed questions you define.
How many questions per request? The verified tests used two and three. Each question adds input and output tokens, so include only what you will act on.
Is there a version alias? The model page lists a jev-latest alias alongside jev-1.13. Pin the explicit version if reproducibility matters.
What does it cost? The model page publishes $0.0462 / M input tokens and $0 / M output tokens. These are vendor figures — confirm current pricing before budgeting.
What is the context window? The model page is inconsistent: 64K in its header, 32K in the provider table. Verify against your own longest input before relying on either.
Scope of this guide
Everything above was verified with three live calls against a single input. It does not cover latency, batching, throughput, or behaviour on inputs unlike the example. The confidence drift is a reproduced observation, not a measured error rate — repeat it on your own data before setting a threshold.
Get started on AiHubMix
jev-1.13 is available through AiHubMix, and the model page carries everything this tutorial referenced: the /v1/systemone endpoint, the three question types with their criteria formats, the full response fields, and the published pricing of $0.0462 / M input tokens with $0 / M output tokens. The jev-latest alias resolves to the newest release, so it stays current as versions move.
Start here: https://aihubmix.com/model/jev-latest
Next steps:
- Create an AiHubMix account and generate an API key
- Copy the request from Step 3 and the
ask()function from Step 4 - Replace
statewith a real record from your own queue - Compare the result against however you classify that record today
If this guide was useful, subscribe for future posts on structured-output models and classification pipelines.



