How to Call Jev on AiHubMix: A Structured Classification Tutorial

AIHubMix5 min read
How to Call Jev on AiHubMix: A Structured Classification Tutorial

Short answer: jev-1.13 does not generate text. You POST a piece of text plus a set of named questions to https://aihubmix.com/v1/systemone, and you get back typed answers keyed by your question names — a category, a score, or a probability. Nothing to parse. Below is a working request, the full response format, and the one design mistake that will quietly cost you accuracy.

What Jev is for

Use it when you need a judgment, not a paragraph: routing a support ticket, scoring severity, flagging urgency, gating content. The model page states it plainly — "No text generation, nothing to parse."

Use a normal chat model instead when you need explanations, summaries, or any free-form output. jev cannot produce those.

Before you start

  • [ ] Python 3 installed (the standard library is enough)
  • [ ] An AiHubMix account and API key
  • [ ] Key exported as AIHUBMIX_API_KEY, never hardcoded in source

You do not need an OpenAI SDK, LangChain, or the requests package.

Step 1 — Know the endpoint

POST https://aihubmix.com/v1/systemone
Authorization: Bearer YOUR_KEY
Content-Type: application/json

This is a model-specific route, not /v1/chat/completions. OpenAI-compatible clients cannot call it. Use plain HTTP.

Step 2 — Choose your question types

There are three, and you can mix them freely in one request:

TypeUse it forYou supply
choicePicking one categorycriteria as a dict: label → definition
scoreRating on an ordered scalecriteria as a list, lowest first
noulA yes/no judgmentnothing beyond instructions

Step 3 — Build the request

state is the text to be judged. questions keys are names you invent; the response uses the same names.

payload = {
    "model": "jev-1.13",
    "state": "Hi, I have been trying to connect my Stripe account for 3 days "
             "and it keeps failing. I am losing sales. Please help ASAP.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle this",
            "criteria": {
                "billing":   "Payment or subscription issues",
                "technical": "Bugs or integration problems",
                "sales":     "Pricing or account questions",
            },
        },
        "frustration": {
            "type": "score",
            "instructions": "How frustrated the customer appears",
            "criteria": [
                "Calm, just stating facts",
                "Frustrated but civil",
                "Very angry, strong language",
            ],
        },
        "is_urgent": {
            "type": "noul",
            "instructions": "The message conveys urgency or time-sensitivity",
        },
    },
}

Step 4 — Send it with the standard library

import json, os, urllib.error, urllib.request

def ask(payload):
    req = urllib.request.Request(
        "https://aihubmix.com/v1/systemone",
        data=json.dumps(payload).encode(),
        headers={"Authorization": "Bearer " + os.environ["AIHUBMIX_API_KEY"],
                 "Content-Type": "application/json"},
        method="POST",
    )
    try:
        with urllib.request.urlopen(req, timeout=60) as r:
            return json.loads(r.read())
    except urllib.error.HTTPError as e:
        raise SystemExit(f"HTTP {e.code}: {e.read().decode(errors='replace')[:500]}")

Keep the request construction inside a function. If you build and send at module level, importing the file re-fires the call and spends tokens.

Step 5 — Read the answers

Each answer stores its value under a key matching its type, so one accessor covers all three:

data = ask(payload)
for name, ans in data["answers"].items():
    kind = ans["type"]
    print(name, kind, ans[kind], ans.get("confidence"))
print("usage:", data.get("usage"))

Verified output:

department   choice  billing   0.51
frustration  score   1         1
is_urgent    noul    1         None
usage: {'input_tokens': 424, 'output_tokens': 73}

Step 6 — Understand the full envelope

The print loop hides useful fields. The raw JSON for a score answer:

{
  "type": "score",
  "score": 1,
  "legend": {"0": "Calm, just stating facts", "1": "Frustrated but civil", "2": "Very angry, strong language"},
  "probabilities": {"0": 0, "1": 1, "2": 0},
  "confidence": 1
}
  • legend maps the index back to your own wording — the integer is self-describing, so you need no lookup table in your code.
  • probabilities gives the full distribution, useful for detecting a near-tie.
  • noul has no confidence field. Its probability is the signal.

The top level also returns usage, id, provider (TypeSafe), and the resolved backend version typesafe/jev-1.13-20260917.

The mistake that costs you accuracy

In the request above, department returned billing with a confidence of 0.51. Repeating the identical request returned the same label at 0.38 — the answer was stable, the stated certainty was not.

The cause is in the criteria, not the model. billing is "Payment or subscription issues" and technical is "Bugs or integration problems." A Stripe integration that keeps failing genuinely matches both. The model was reporting an ambiguity that was written into the schema.

Two rules before production:

  1. Write choice criteria that are mutually exclusive. If a human would hesitate between two labels, the model will too.
  2. Set a confidence threshold and route low-confidence results to a human queue instead of accepting them as decisions. A model that admits uncertainty is more valuable than one that hides it.

Pre-launch checklist

  • [ ] Key loaded from environment or a 600-mode file your VCS ignores
  • [ ] choice criteria reviewed for overlap
  • [ ] Confidence threshold defined, with a fallback path below it
  • [ ] noul answers handled separately — they carry no confidence
  • [ ] Unused questions removed (3 questions cost 424 input tokens; 2 cost 355)
  • [ ] Non-2xx responses logged with the raw body

FAQ

Can I use the OpenAI SDK? No. /v1/systemone is not an OpenAI-compatible route.

Can jev return a sentence or summary? No. It answers only the typed questions you define.

How many questions per request? The verified tests used two and three. Each question adds input and output tokens, so include only what you will act on.

Is there a version alias? The model page lists a jev-latest alias alongside jev-1.13. Pin the explicit version if reproducibility matters.

What does it cost? The model page publishes $0.0462 / M input tokens and $0 / M output tokens. These are vendor figures — confirm current pricing before budgeting.

What is the context window? The model page is inconsistent: 64K in its header, 32K in the provider table. Verify against your own longest input before relying on either.

Scope of this guide

Everything above was verified with three live calls against a single input. It does not cover latency, batching, throughput, or behaviour on inputs unlike the example. The confidence drift is a reproduced observation, not a measured error rate — repeat it on your own data before setting a threshold.

Get started on AiHubMix

jev-1.13 is available through AiHubMix, and the model page carries everything this tutorial referenced: the /v1/systemone endpoint, the three question types with their criteria formats, the full response fields, and the published pricing of $0.0462 / M input tokens with $0 / M output tokens. The jev-latest alias resolves to the newest release, so it stays current as versions move.

Start here: https://aihubmix.com/model/jev-latest

Next steps:

  1. Create an AiHubMix account and generate an API key
  2. Copy the request from Step 3 and the ask() function from Step 4
  3. Replace state with a real record from your own queue
  4. Compare the result against however you classify that record today

If this guide was useful, subscribe for future posts on structured-output models and classification pipelines.