Overview
Probabilities. Without the prose.
Fez returns structured probabilities without generating a text answer. Applications can load a selected checkpoint directly.
PUBLIC BENCHMARK CANDIDATE
63475109e543fe5f7fd2698ec27f6ac0b828c67074606409ad71697abca30da3Model quality
Accuracy tied; confidence quality regressed.
FEZ ACCURACY
63.64%
147 of 231 public items correct
FEZ BRIER LOSS
0.463608
Probability error. Lower is better.
FEZ CONFIDENT MISTAKES
8
Wrong at ≥90% confidence.
Fez corrected 6 reference-model errors and introduced 6. No general improvement is established.
JevBench public comparison
231 released public items: 48 easy, 72 original, 111 hard
| Metric | Fez candidateExperimental checkpoint | Published Kev 0.8BComparison baseline |
|---|---|---|
| AccuracyHigher is better | 63.64% | 63.64% |
| Correct answersSame public items | 147 / 231 | 147 / 231 |
| Brier lossLower is better | 0.463608 | 0.451648 |
| Calibration error (ECE)Lower is better | 0.113615 | 0.044050 |
| Confident mistakesWrong at ≥90% confidence | 8 | 3 |
| Median latencyLocalhost HTTP | 42.08 ms | 52.78 ms |
| p95 latencyLocalhost HTTP | 205.26 ms | 209.36 ms |
| Valid probability outputs | 100.00% | 100.00% |
Brier loss measures probability error; ECE measures the gap between confidence and observed accuracy. Lower is better for both. Timing is one serial pass per model; Kev then Fez, on NVIDIA GeForce RTX 4090, via localhost HTTP. This does not establish a reliable speedup or production SLA.
Accuracy by difficulty
| Subset | Fez | Kev baseline |
|---|---|---|
| easy | 48 / 48100.00% | 48 / 48100.00% |
| original | 57 / 7279.17% | 58 / 7280.56% |
| hard | 42 / 11137.84% | 41 / 11136.94% |
- Data mode
- recorded_experiment
- Completed, recorded experiment
- Verified at
- Evidence time; not page-refresh time
- Hardware / precision
- NVIDIA GeForce RTX 4090
- FP32. Same runtime for both models.
Method, checkpoint identities & limitations
Fez fine-tunes a published Kev 0.8B checkpoint built on Qwen3.5-0.8B-Base. The unchanged published checkpoint is the baseline in this comparison.
No official JevBench rank or composite score is available: private and sealed tests were not run. Hosted cost was not measured. Other experiment suites are documented separately and cannot form an improvement line with this comparison.
- Experiment / collection started
- jevbench-public-001 /
- Fez checkpoint SHA-256
63475109e543fe5f7fd2698ec27f6ac0b828c67074606409ad71697abca30da3- Published Kev checkpoint SHA-256
af9331c6c6331099ce2d28a0298620ca540c991f781b6167ceb0f177f235190f- Public dataset SHA-256
dc3995d8ae1e2fc8e81ce38431add509eb8bb39b85aadfd0c7c32079382dde51- Harness revision
26eb72d4e0e60d8ace0adfc77a384063442561cd- Latency scope
- localhost HTTP including encoding, inference and serialization; excludes model loading and three synthetic warmup calls.
- Saved temperatures
- Fez: 1.148698. Kev: 2.351096.
- Public subset only; no official leaderboard rank or composite.
- Both saved temperatures retained; no calibration or training on JevBench.
- No exact normalized state overlap found in recorded local training; upstream contamination and semantic similarity remain unknown.
- Single serial timing pass, not a repeated or production-speed study.
- Hosted cost unmeasured; no cost estimate is presented.
The experimental Fez checkpoint is not distributed, so the full comparison cannot be reproduced from the checkout alone.
Training & evaluation
Train. Submit. Evaluate. Reward.
- 01
Train
Miners train compatible adapters and decision heads with a fixed one-epoch recipe and different seeds.
- 02
Submit
Each miner freezes a candidate and signs its checkpoint hash for the validator.
- 03
Evaluate
The validator verifies the checkpoint and runs its own evaluation of the model’s probabilities.
- 04
Reward
Family-macro Brier skill determines proposed weights across eligible candidates.
One testnet round completed
Training, signed submissions, evaluation, and revealed chain weights are recorded below. A live submission queue and round feed are not connected.
View the verified roundThe local loop works on Apple Silicon and an RTX 4090. Accuracy and latency are diagnostics, not separate reward components. The subnet’s family-macro Brier and JevBench Brier use different aggregation rules.
Evaluation contractParticipants
Three miners. One validator. One host.
Three registered miners and one validator completed the recorded subnet 579 rehearsal.
- Miners · UIDs 1–3
- Trained and submitted three distinct checkpoints
- RECORDED PARTICIPATION
- Validator · UID 0
- Evaluated the checkpoints and published weights
- RECORDED PARTICIPATION
One operator ran all four processes on one host. Services exited after the round; these counts do not indicate current availability or independent operators.
Miner guideTestnet publication
Bittensor testnet · Subnet 579
First training-to-chain round completed. Revealed weights verified.
Round evidence (JSON) ↗- Network
- Bittensor testnet
- Publication receipt
- 8081301-0008
- Verified at block
- 8,081,345
- Round status
- Completed · recorded rehearsal
| Miner UID | Correct / 224 | Requested weight | On-chain value |
|---|---|---|---|
| 1 | 183 / 224 | 34.09% | 65,535 |
| 2 | 180 / 224 | 33.60% | 64,598 |
| 3 | 180 / 224 | 32.31% | 62,125 |
Weights are based on family-macro Brier skill. The chain stores maximum-scaled integers; their normalized proportions matched the requested allocation within quantization tolerance. Verified . This is the observation time, not a live refresh.
Each 0.8B checkpoint trained for one epoch on 224 examples, calibrated on 112 questions, and was evaluated on the same 224 test questions. This reused synthetic development benchmark is separate from the public JevBench comparison above. All compute ran on one Apple M4 Pro with 24 GiB memory, with GPU jobs serialized.
This closed rehearsal demonstrates the training-to-chain path. It does not establish mainnet deployment, miner earnings, open competition, or automatic model promotion.
Fez decision model. Public evidence, read-only access.