Research · Open about the evidence

Decision model research.

Fez is an experimental 0.8B decision model for yes/no questions, choices, and scores. It returns structured probabilities without generating prose.

Experimental candidate. No public model release or automatic winner promotion yet.

BITTENSOR TESTNET · SUBNET 579First training round verified on-chain

Overview

Probabilities. Without the prose.

Fez returns structured probabilities without generating a text answer. Applications can load a selected checkpoint directly.

PUBLIC BENCHMARK CANDIDATE

63475109e543fe5f7fd2698ec27f6ac0b828c67074606409ad71697abca30da3

Model quality

Accuracy tied; confidence quality regressed.

FEZ ACCURACY

63.64%

147 of 231 public items correct

FEZ BRIER LOSS

0.463608

Probability error. Lower is better.

FEZ CONFIDENT MISTAKES

8

Wrong at ≥90% confidence.

Fez corrected 6 reference-model errors and introduced 6. No general improvement is established.

JevBench public comparison

231 released public items: 48 easy, 72 original, 111 hard

Source data (JSON)
Recorded Fez and published Kev results on the same public items
MetricFez candidateExperimental checkpointPublished Kev 0.8BComparison baseline
AccuracyHigher is better63.64%63.64%
Correct answersSame public items147 / 231147 / 231
Brier lossLower is better0.4636080.451648
Calibration error (ECE)Lower is better0.1136150.044050
Confident mistakesWrong at ≥90% confidence83
Median latencyLocalhost HTTP42.08 ms52.78 ms
p95 latencyLocalhost HTTP205.26 ms209.36 ms
Valid probability outputs100.00%100.00%

Brier loss measures probability error; ECE measures the gap between confidence and observed accuracy. Lower is better for both. Timing is one serial pass per model; Kev then Fez, on NVIDIA GeForce RTX 4090, via localhost HTTP. This does not establish a reliable speedup or production SLA.

Accuracy by difficulty

Correct and total counts, with accuracy, for each public subset
SubsetFezKev baseline
easy48 / 48100.00%48 / 48100.00%
original57 / 7279.17%58 / 7280.56%
hard42 / 11137.84%41 / 11136.94%
Data mode
recorded_experiment
Completed, recorded experiment
Verified at
Evidence time; not page-refresh time
Hardware / precision
NVIDIA GeForce RTX 4090
FP32. Same runtime for both models.
Method, checkpoint identities & limitations

Fez fine-tunes a published Kev 0.8B checkpoint built on Qwen3.5-0.8B-Base. The unchanged published checkpoint is the baseline in this comparison.

No official JevBench rank or composite score is available: private and sealed tests were not run. Hosted cost was not measured. Other experiment suites are documented separately and cannot form an improvement line with this comparison.

Experiment / collection started
jevbench-public-001 /
Fez checkpoint SHA-256
63475109e543fe5f7fd2698ec27f6ac0b828c67074606409ad71697abca30da3
Published Kev checkpoint SHA-256
af9331c6c6331099ce2d28a0298620ca540c991f781b6167ceb0f177f235190f
Public dataset SHA-256
dc3995d8ae1e2fc8e81ce38431add509eb8bb39b85aadfd0c7c32079382dde51
Harness revision
26eb72d4e0e60d8ace0adfc77a384063442561cd
Latency scope
localhost HTTP including encoding, inference and serialization; excludes model loading and three synthetic warmup calls.
Saved temperatures
Fez: 1.148698. Kev: 2.351096.
  • Public subset only; no official leaderboard rank or composite.
  • Both saved temperatures retained; no calibration or training on JevBench.
  • No exact normalized state overlap found in recorded local training; upstream contamination and semantic similarity remain unknown.
  • Single serial timing pass, not a repeated or production-speed study.
  • Hosted cost unmeasured; no cost estimate is presented.

The experimental Fez checkpoint is not distributed, so the full comparison cannot be reproduced from the checkout alone.

Training & evaluation

Train. Submit. Evaluate. Reward.

  1. 01

    Train

    Miners train compatible adapters and decision heads with a fixed one-epoch recipe and different seeds.

  2. 02

    Submit

    Each miner freezes a candidate and signs its checkpoint hash for the validator.

  3. 03

    Evaluate

    The validator verifies the checkpoint and runs its own evaluation of the model’s probabilities.

  4. 04

    Reward

    Family-macro Brier skill determines proposed weights across eligible candidates.

One testnet round completed

Training, signed submissions, evaluation, and revealed chain weights are recorded below. A live submission queue and round feed are not connected.

View the verified round

The local loop works on Apple Silicon and an RTX 4090. Accuracy and latency are diagnostics, not separate reward components. The subnet’s family-macro Brier and JevBench Brier use different aggregation rules.

Evaluation contract

Participants

Three miners. One validator. One host.

Three registered miners and one validator completed the recorded subnet 579 rehearsal.

Miners · UIDs 1–3
Trained and submitted three distinct checkpoints
RECORDED PARTICIPATION
Validator · UID 0
Evaluated the checkpoints and published weights
RECORDED PARTICIPATION

One operator ran all four processes on one host. Services exited after the round; these counts do not indicate current availability or independent operators.

Miner guide

Testnet publication

Bittensor testnet · Subnet 579

First training-to-chain round completed. Revealed weights verified.

Round evidence (JSON) ↗
Network
Bittensor testnet
Publication receipt
8081301-0008
Verified at block
8,081,345
Round status
Completed · recorded rehearsal
Three evaluated miners and their verified testnet weight allocation
Miner UIDCorrect / 224Requested weightOn-chain value
1183 / 22434.09%65,535
2180 / 22433.60%64,598
3180 / 22432.31%62,125

Weights are based on family-macro Brier skill. The chain stores maximum-scaled integers; their normalized proportions matched the requested allocation within quantization tolerance. Verified . This is the observation time, not a live refresh.

Each 0.8B checkpoint trained for one epoch on 224 examples, calibrated on 112 questions, and was evaluated on the same 224 test questions. This reused synthetic development benchmark is separate from the public JevBench comparison above. All compute ran on one Apple M4 Pro with 24 GiB memory, with GPU jobs serialized.

This closed rehearsal demonstrates the training-to-chain path. It does not establish mainnet deployment, miner earnings, open competition, or automatic model promotion.

Fez decision model. Public evidence, read-only access.