THX-01: Open-Sourcing a 322M Non-Autoregressive Decision Model Matching Claude 3.5 Sonnet Speed and Accuracy
- Authors

- Name
- Nino
- Occupation
- Senior Tech Editor
Using general-purpose autoregressive Large Language Models (LLMs) like Claude 3.5 Sonnet or GPT-4o for deterministic tasks like support ticket classification, routing, and intent extraction often introduces unnecessary latency, high inference costs, and brittle JSON parsing logic. While frontier foundation models excel at open-ended generation, standard production classification tasks demand sub-50ms responses, high batch throughput, strictly calibrated probability distributions, and 100% schema reliability.
To bridge this gap, THX-01 has been released as an open-source (Apache 2.0) specialized decision model. With only 322M parameters, THX-01 matches frontier LLM accuracy on complex multi-lingual ticket routing benchmarks while executing inference in approximately 10 ms on a single GPU and scaling to over 3,000 decisions per second in batch mode.
In this article, we analyze the non-autoregressive architecture of THX-01, explore benchmark comparisons against leading models, walk through practical Python implementations, and demonstrate how to integrate THX-01 alongside multi-LLM orchestration platforms like n1n.ai for optimal speed and reliability.
Architecture Breakdown: Non-Autoregressive Typed Decisions
Traditional LLM pipelines for classification rely on token-by-token text generation followed by structured extraction (e.g., JSON schema validation). This introduces several architectural bottlenecks:
- Token Generation Overhead: Generating 20 to 50 tokens requires dozens of autoregressive forward passes through billions of parameters.
- Parsing Failure Modes: Syntax errors, missing keys, or unexpected markdown formatting require retry loops and validation layer logic.
- Uncalibrated Confidence: Raw softmax logits from text generation outputs rarely translate to true statistical probabilities, making confidence thresholding unreliable.
THX-01 eliminates these constraints through a dual-component architecture consisting of an mmBERT-base encoder combined with a specialized 15M decision head.
+-------------------------------------------------------------------------+
| THX-01 Architecture |
+-------------------------------------------------------------------------+
| Input State (Text / Context) + Typed Questions (Schema definition) |
+-------------------------------------------------------------------------+
|
v
+-------------------------------------------------------------------------+
| mmBERT-Base Encoder (307M Parameters) |
+-------------------------------------------------------------------------+
|
v
+-------------------------------------------------------------------------+
| Typed Decision Head (15M Parameters) |
+-------------------------------------------------------------------------+
|
+---------------------------+---------------------------+
| | |
v v v
+------------------+ +------------------+ +------------------+
| Choice (Categorical) | Bool (Yes/No) | Excerpt & Number|
| Probabilities | | Probabilities | | Extraction |
+------------------+ +------------------+ +------------------+
Core Technical Features
- Single Forward Pass Execution: THX-01 is fully non-autoregressive. It evaluates input state and typed question schemas simultaneously in a single forward pass, returning exact probability values for every choice without token generation.
- Built-in Schema Validation: Eliminates JSON parsing, regex matching, or retry loops entirely. The model natively outputs typed structures.
- Support for Multi-Type Questions: Supports
choice(categorical distributions),bool(binary decisions),score(continuous ratings),number(numeric value extraction from context),excerpt(exact span extraction with token offsets), and optionalcite: trueflags for verifiable provenance. - Strict Calibration (ECE 0.003): Trained under strictly proper scoring-rule rewards, THX-01 achieves an Expected Calibration Error (ECE) of 0.003. A predicted probability of 98.4% statistically guarantees an empirical accuracy of ~98.4%.
- Native Wire Format: Out of the box, THX-01 speaks the
/v1/systemonewire protocol, serving as a drop-in replacement for enterprise TypeSafe microservices.
Benchmark Performance Analysis
To evaluate production reliability under realistic conditions, THX-01 was benchmarked against leading open and proprietary models across a dataset of 2,843 multi-lingual customer support tickets spanning 15 classification categories in Azerbaijani (AZ), Russian (RU), English (EN), and Turkish (TR).
| Model | Average Accuracy (%) | Single-Request Latency | Execution Engine / Size |
|---|---|---|---|
| THX-01 | 98.4% | ~10 ms | 322M (PyTorch / On-Premise GPU or CPU) |
| Claude 3.5 Sonnet | 98.5% | 1,500 ms | Proprietary API Cloud |
| Wahoo 1.5 | 97.8% | 145 ms | Specialized Mid-tier Model |
| TypeSafe Jev 1.13 | 97.4% | 331 ms | Enterprise Micro-Model |
| Kev-4B | 92.8% | 830 ms | 4B Autoregressive Model |
Key Takeaways from Benchmarks
- 150x Latency Reduction: THX-01 delivers classification performance identical to Claude 3.5 Sonnet (98.4% vs 98.5%) while completing requests in ~10 ms compared to 1,500 ms—a 150x speedup.
- Throughput Scaling: Because of its compact 322M parameter footprint, a single mid-range enterprise GPU can process ~3,000 decisions per second via batched inference. Furthermore, THX-01 runs efficiently on standard modern CPUs without specialized hardware acceleration.
- Superior Calibration Over General LLMs: Unlike general LLMs that exhibit overconfidence on out-of-domain edge cases, THX-01's calibrated low ECE allows engineers to establish reliable confidence thresholds for automated routing.
Step-by-Step Python Implementation
Installing THX-01 takes seconds via PyPI:
pip install thx01
Basic Multi-Choice Classification Example
The following code demonstrates how to load THX-01 from HuggingFace (doofz/THX-01) and run categorical classification on incoming user requests:
import thx01
# Initialize and load the model weights
agent = thx01.load("doofz/THX-01")
# Define input state
user_ticket = "My credit card was billed twice for order #84920. Please refund the extra charge."
# Submit typed question schema
decisions = agent.decide(
state=user_ticket,
schema=\{
"target_department": \{
"type": "choice