Complete protocol-screened leaderboard
Educational scores retain the 20 August historical snapshot. High-Motion v2 compares CPU-rescored archived baselines with one new GRT run; no baseline inference was repeated. Download CSV · Versioned source hashes and comparison limits · Public reproduction code · Corrected-reference reproduction guide.
29 Educational results, 19 corrected High-Motion preview results, and 12 educational GRT comparison rows (60 CSV records). Educational: 634 QA / 317 videos. High-Motion v2: the fixed 1,000-record preview of 3,243 source records; metric-specific valid-reference coverage is shown separately. All numeric cells retain full precision in their title and downloadable data.
Educational High-FPS Videos
29 published methods. Rank by reported Open MOS, then Token F1; missing Open MOS sorts last and is never filled in. Open MOS judge: Qwen/Qwen3-VL-32B-Instruct. The verified GRT profiles use eight sampled frames, not a measured high-FPS operating point.
| Rank | Model | Method ID | Items | Open MOS ↑ | Token F1 ↑ | CER ↓ | WER ↓ | Exact match ↑ | Patch recompute ↓ | Reference patch compute ↓ | Sampling density (fps) | Mean throughput (fps) ↑ | Source |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro Preview (gemini-3.1-pro-preview) | gemini-3.1-pro-preview | 634 | 1.6041 | 0.159633 | 1.26881 | 1.24166 | 0 | — | — | — | — | gemini |
| 2 | GRT (Qwen2.5-VL 7B, route floor 0.80 / 0.55, cap 48) | grt_qwen2_5_vl_7b_dual_floor_s080_o055_cap48 | 634 | 1.59148 | 0.0489087 | 1.06401 | 1.05377 | 0 | 0.847719 | 0.847719 | 0.00747917 | 2.48043 | open |
| 3 | GRT (Qwen2.5-VL 3B, threshold 0.3) | grt_qwen2_5_vl_3b_t03 | 634 | 1.58991 | 0.0995837 | 1.13915 | 1.08351 | 0 | 0.885659 | 0.885659 | 0.00747917 | 1.67264 | open |
| 4 | Qwen3-VL 8B Instruct | qwen3_vl_8b | 634 | 1.29811 | 0.0360398 | 1.05889 | 1.05198 | 0 | — | — | — | — | open |
| 5 | Qwen2.5-VL 7B Instruct | qwen2_5_vl_7b | 634 | 1.28391 | 0.0347846 | 1.04624 | 1.03975 | 0 | — | — | — | — | open |
| 6 | Gemini 3.6 Flash (gemini-3.6-flash) | gemini-3.6-flash | 634 | 1.24921 | 0.0814611 | 1.10168 | 1.06796 | 0 | — | — | — | — | gemini |
| 7 | Qwen3-VL 2B Instruct | qwen3_vl_2b | 634 | 1.14511 | 0.0331448 | 1.0455 | 1.0365 | 0 | — | — | — | — | open |
| 8 | Qwen3-VL 4B Instruct | qwen3_vl_4b | 634 | 1.1388 | 0.0348606 | 1.05593 | 1.05119 | 0 | — | — | — | — | open |
| 9 | Qwen2.5-VL 3B Instruct | qwen2_5_vl_3b | 634 | 1.13565 | 0.0343388 | 1.04393 | 1.03399 | 0 | — | — | — | — | open |
| 10 | Qwen2.5-VL 72B Instruct | qwen2_5_vl_72b | 634 | 0.944795 | 0.0268628 | 1.05934 | 1.04789 | 0 | — | — | — | — | open |
| 11 | Qwen2-VL 2B Instruct | qwen2_vl_2b | 634 | 0.790221 | 0.0200933 | 1.03692 | 1.0334 | 0 | — | — | — | — | open |
| 12 | Qwen2.5-VL 32B Instruct | qwen2_5_vl_32b | 634 | 0.5 | 0.0250276 | 1.05448 | 1.04856 | 0 | — | — | — | — | open |
| 13 | Qwen3-VL 32B Instruct | qwen3_vl_32b | 634 | 0.175079 | 0.0184676 | 1.05239 | 1.05085 | 0 | — | — | — | — | open |
| 14 | GRT (LLaVA-OneVision Qwen2 0.5B, verified Route31) | grt_llava_onevision_0_5b_hf_route31_t0001 | 634 | 0.119874 | 0.0141964 | 1.11093 | 1.10273 | 0 | 0.867154 | 0.867154 | 0.00747917 | 4.86514 | open |
| 15 | LLaVA-OneVision Qwen2 0.5B | llava_onevision_0_5b | 634 | 0.116719 | 0.0140879 | 1.03765 | 1.03756 | 0 | — | — | — | — | open |
| 16 | LLaVA-OneVision 1.5 8B Instruct | llava_onevision_1_5_8b | 634 | — | 0.0349748 | 1.07243 | 1.05686 | 0 | — | — | — | — | open |
| 17 | Qwen2-VL 7B Instruct | qwen2_vl_7b | 634 | — | 0.0305096 | 1.05309 | 1.04788 | 0 | — | — | — | — | open |
| 18 | Phi-4 Multimodal Instruct | phi4_multimodal | 634 | — | 0.0275287 | 1.03008 | 1.02762 | 0 | — | — | — | — | open |
| 19 | InternVL3 1B | internvl3_1b | 634 | — | 0.0271622 | 1.062 | 1.0606 | 0 | — | — | — | — | open |
| 20 | InternVL3 8B | internvl3_8b | 634 | — | 0.0250303 | 1.07046 | 1.06601 | 0 | — | — | — | — | open |
| 21 | VideoLLaMA3 7B | videollama3_7b | 634 | — | 0.0249369 | 1.05106 | 1.04975 | 0 | — | — | — | — | open |
| 22 | LongVA 7B | longva_7b | 634 | — | 0.0246616 | 1.04678 | 1.04241 | 0 | — | — | — | — | open |
| 23 | InternVL3 2B | internvl3_2b | 634 | — | 0.0223091 | 1.03408 | 1.03574 | 0 | — | — | — | — | open |
| 24 | InternVL2.5 8B | internvl2_5_8b | 634 | — | 0.0188533 | 1.05738 | 1.05719 | 0 | — | — | — | — | open |
| 25 | InternVL2.5 4B | internvl2_5_4b | 634 | — | 0.01792 | 1.02051 | 1.02362 | 0 | — | — | — | — | open |
| 26 | InternVL2.5 1B | internvl2_5_1b | 634 | — | 0.0178427 | 1.03147 | 1.0336 | 0 | — | — | — | — | open |
| 27 | LLaVA-OneVision Qwen2 7B | llava_onevision_original | 634 | — | 0.0175202 | 1.04056 | 1.03982 | 0 | — | — | — | — | open |
| 28 | VideoLLaMA3 2B | videollama3_2b | 634 | — | 0.0158476 | 1.0245 | 1.02683 | 0 | — | — | — | — | open |
| 29 | InternVL2.5 2B | internvl2_5_2b | 634 | — | 0.0119298 | 1.01729 | 1.02037 | 0 | — | — | — | — | open |
High-Motion v2: right-ring reference correction, preview-1000
19 methods: 18 archived baselines rescored on CPU and one new HF 0.5B GRT run. All use the same fixed first 1,000 source records, not a full 3,243-record evaluation. The reference target is rightRingFingerMetacarpal, the named right-palm/ring-finger-base proxy; questions and original sampled positions are unchanged. This versioned correction does not claim to recover the original annotation constructor.
Rank by Grid Accuracy descending; exact ties share a competition rank and are ordered by method ID. Null scores are unranked. Every metric is a macro mean over its defined per-record values, with visible metric-specific row and slot/edge coverage. A dash means undefined under the reference mask, not zero. Invalid slots never shift later predictions; FDE uses the original final slot, and transitions require adjacent valid original slots. Token F1 uses canonical label bags on valid positions, with surplus outputs penalized. All 1,000 records remain counted even when no positions are scoreable.
New GRT predictions versus rescored, configuration-checked archived predictions; archived baseline weight revisions and consumed-tensor identity are unproven. This is not a freshly rerun byte-identical paired experiment or a statistical-significance claim.
| Rank | Model | Method ID | Cached/new records | Grid Accuracy ↑ | Grid Accuracy ↑ coverage | Grid ADE ↓ | Grid ADE ↓ coverage | Grid FDE ↓ | Grid FDE ↓ coverage | Transition Accuracy ↑ | Transition Accuracy ↑ coverage | Token F1 ↑ | Token F1 ↑ coverage | Prediction provenance |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | LLaVA-OneVision-2-8B-Instruct | llava_onevision_2_8b | 1000 | 0.512279 | 861 rows / 6015 slots | 0.279313 | 861 rows / 6015 slots | 0.206212 | 669 rows / 669 slots | 0.731701 | 852 rows / 5101 edges | 0.513393 | 861 rows / 6015 slots | archived_baseline |
| 2 | Qwen3-VL-8B-Instruct | qwen3_vl_8b | 1000 | 0.426745 | 861 rows / 6015 slots | 0.357624 | 861 rows / 6015 slots | 0.310443 | 669 rows / 669 slots | 0.734879 | 852 rows / 5101 edges | 0.426938 | 861 rows / 6015 slots | archived_baseline |
| 3 | LLaVA-OneVision-1.5-8B-Instruct | llava_onevision_1_5_8b | 1000 | 0.41958 | 861 rows / 6015 slots | 0.374625 | 861 rows / 6015 slots | 0.35337 | 669 rows / 669 slots | 0.729748 | 852 rows / 5101 edges | 0.41958 | 861 rows / 6015 slots | archived_baseline |
| 4 | VideoLLaMA3-7B | videollama3_7b | 1000 | 0.416147 | 861 rows / 6015 slots | 0.479898 | 861 rows / 6015 slots | 0.591313 | 669 rows / 669 slots | 0.440789 | 852 rows / 5101 edges | 0.462768 | 861 rows / 6015 slots | archived_baseline |
| 5 | Qwen3-VL-32B-Instruct | qwen3_vl_32b | 1000 | 0.360901 | 861 rows / 6015 slots | 0.384434 | 861 rows / 6015 slots | 0.392697 | 669 rows / 669 slots | 0.72526 | 852 rows / 5101 edges | 0.363749 | 861 rows / 6015 slots | archived_baseline |
| 6 | Qwen3-VL-4B-Instruct | qwen3_vl_4b | 1000 | 0.359986 | 861 rows / 6015 slots | 0.359877 | 861 rows / 6015 slots | 0.368279 | 669 rows / 669 slots | 0.735113 | 852 rows / 5101 edges | 0.359986 | 861 rows / 6015 slots | archived_baseline |
| 7 | Qwen3-VL-2B-Instruct | qwen3_vl_2b | 1000 | 0.275784 | 861 rows / 6015 slots | 0.423396 | 861 rows / 6015 slots | 0.471795 | 669 rows / 669 slots | 0.313081 | 852 rows / 5101 edges | 0.35428 | 861 rows / 6015 slots | archived_baseline |
| 8 | Qwen2.5-VL-72B-Instruct | qwen2_5_vl_72b | 1000 | 0.235602 | 861 rows / 6015 slots | 0.513407 | 861 rows / 6015 slots | 0.569024 | 669 rows / 669 slots | 0.648879 | 852 rows / 5101 edges | 0.258947 | 861 rows / 6015 slots | archived_baseline |
| 9 | Qwen2.5-VL-32B-Instruct | qwen2_5_vl_32b | 1000 | 0.176484 | 861 rows / 6015 slots | 0.505219 | 861 rows / 6015 slots | 0.504983 | 669 rows / 669 slots | 0.56006 | 852 rows / 5101 edges | 0.189601 | 861 rows / 6015 slots | archived_baseline |
| 10 | LongVA-7B | longva_7b | 1000 | 0.0760605 | 861 rows / 6015 slots | 0.770476 | 861 rows / 6015 slots | 0.632581 | 669 rows / 669 slots | 0.120532 | 852 rows / 5101 edges | 0.155908 | 861 rows / 6015 slots | archived_baseline |
| 11 | Phi-4-multimodal-instruct | phi4_multimodal | 1000 | 0.0722015 | 861 rows / 6015 slots | 1.1359 | 861 rows / 6015 slots | 1.21451 | 669 rows / 669 slots | 0.0180751 | 852 rows / 5101 edges | 0.11382 | 861 rows / 6015 slots | archived_baseline |
| 12 | Qwen2-VL-2B-Instruct | qwen2_vl_2b | 1000 | 0.0656546 | 861 rows / 6015 slots | 0.771788 | 861 rows / 6015 slots | 0.625801 | 669 rows / 669 slots | 0.225872 | 852 rows / 5101 edges | 0.105367 | 861 rows / 6015 slots | archived_baseline |
| 13 | LLaVA-OneVision HF 7B | llava_onevision_original | 1000 | 0.0624772 | 861 rows / 6015 slots | 0.881176 | 861 rows / 6015 slots | 0.846691 | 669 rows / 669 slots | 0.185616 | 852 rows / 5101 edges | 0.116204 | 861 rows / 6015 slots | archived_baseline |
| 14 | VideoLLaMA3-2B | videollama3_2b | 1000 | 0.0549375 | 861 rows / 6015 slots | 1.12491 | 861 rows / 6015 slots | 1.15021 | 669 rows / 669 slots | 0.000503018 | 852 rows / 5101 edges | 0.115104 | 861 rows / 6015 slots | archived_baseline |
| 15 | Qwen2-VL-7B-Instruct | qwen2_vl_7b | 1000 | 0.0524114 | 861 rows / 6015 slots | 0.855721 | 861 rows / 6015 slots | 0.905482 | 669 rows / 669 slots | 0.30432 | 852 rows / 5101 edges | 0.069898 | 861 rows / 6015 slots | archived_baseline |
| 16 | GRT · LLaVA-OneVision HF 0.5B (motion SSIM 0.001) | grt_llava_hf_0_5b_motion_ssim_t0001 | 1000 | 0.0495036 | 861 rows / 6015 slots | 1.0862 | 861 rows / 6015 slots | 1.06721 | 669 rows / 669 slots | 0.00534038 | 852 rows / 5101 edges | 0.0450525 | 861 rows / 6015 slots | new_grt |
| 17 | LLaVA-OneVision HF 0.5B | llava_onevision_0_5b | 1000 | 0.0468683 | 861 rows / 6015 slots | 1.09641 | 861 rows / 6015 slots | 1.07772 | 669 rows / 669 slots | 0.0056841 | 852 rows / 5101 edges | 0.0418983 | 861 rows / 6015 slots | archived_baseline |
| 18 | Qwen2.5-VL-3B-Instruct | qwen2_5_vl_3b | 1000 | 0.0150863 | 861 rows / 6015 slots | 0.932863 | 861 rows / 6015 slots | 0.998825 | 669 rows / 669 slots | 0.576009 | 852 rows / 5101 edges | 0.0201869 | 861 rows / 6015 slots | archived_baseline |
| 19 | Qwen2.5-VL-7B-Instruct | qwen2_5_vl_7b | 1000 | 0.0149204 | 861 rows / 6015 slots | 1.18876 | 861 rows / 6015 slots | 1.30077 | 669 rows / 669 slots | 0.579396 | 852 rows / 5101 edges | 0.0135712 | 861 rows / 6015 slots | archived_baseline |
GRT versus its corresponding HF 0.5B baseline
GRT exceeds the corresponding HF 0.5B baseline on observed Grid Accuracy. Point tolerance: 1e-12. Every observed difference is retained, including regressions. This is a point-estimate comparison, not a statistical-significance claim.
| Metric | Rescored HF 0.5B baseline | GRT | GRT minus baseline | Oriented improvement (positive is better) | Exceeds point tolerance |
|---|---|---|---|---|---|
| Grid Accuracy ↑ | 0.0468683 | 0.0495036 | 0.00263536 | 0.00263536 | yes |
| Grid ADE ↓ | 1.09641 | 1.0862 | -0.0102074 | 0.0102074 | yes |
| Grid FDE ↓ | 1.07772 | 1.06721 | -0.010501 | 0.010501 | yes |
| Transition Accuracy ↑ | 0.0056841 | 0.00534038 | -0.000343729 | -0.000343729 | no |
| Token F1 ↑ | 0.0418983 | 0.0450525 | 0.00315414 | 0.00315414 | yes |
Legacy High-Motion results remain withheld
The old target/reference hold and immutable 27-run historical protocol audit remain intact. No old High-Motion score is mixed into the corrected v2 table or its CSV cohort. Reference, scorer, prediction and release provenance are available in the versioned numeric audit.
GRT vs archived and matched controls
All three promoted educational candidates exceed every contracted Open MOS and Token F1 floor, with 11.43–15.23% fewer patch projections than the matched all-patch route. LLaVA-OneVision 7B failed its MOS gate and is not promoted. These are observed point estimates, not statistical-significance claims. Archived Qwen baselines differ from stronger matched controls: their entire gap must not be attributed to GRT.
| Family | Control | Method ID | Items | Open MOS ↑ | Token F1 ↑ | Patch recompute ↓ | Mean throughput (fps) ↑ | Mean request time (s) ↓ |
|---|---|---|---|---|---|---|---|---|
| LLaVA-OneVision 0.5B Route31 | Archived public baseline | llava_onevision_0_5b | 634 | 0.116719 | 0.0140879 | — | — | — |
| LLaVA-OneVision 0.5B Route31 | Matched quality baseline | llava_onevision_0_5b_hf_route31_base | 634 | 0.116719 | 0.0139815 | 1 | 4.81928 | 2.0401 |
| LLaVA-OneVision 0.5B Route31 | Matched all-patch control | llava_onevision_0_5b_hf_route31_exact | 634 | 0.116719 | 0.0139815 | 1 | 4.8239 | 2.00908 |
| LLaVA-OneVision 0.5B Route31 | GRT candidate | grt_llava_onevision_0_5b_hf_route31_t0001 | 634 | 0.119874 | 0.0141964 | 0.867154 | 4.86514 | 2.00911 |
| Qwen2.5-VL 3B t03 | Archived public baseline | qwen2_5_vl_3b | 634 | 1.13565 | 0.0343388 | — | — | — |
| Qwen2.5-VL 3B t03 | Matched all-patch control | qwen2_5_vl_3b_grt_all | 634 | 1.55836 | 0.0991358 | 1 | 1.74994 | 5.42461 |
| Qwen2.5-VL 3B t03 | Matched quality baseline | qwen2_5_vl_3b_quality_base | 634 | 1.55363 | 0.0988916 | 1 | 1.63172 | 5.88014 |
| Qwen2.5-VL 3B t03 | GRT candidate | grt_qwen2_5_vl_3b_t03 | 634 | 1.58991 | 0.0995837 | 0.885659 | 1.67264 | 5.70157 |
| Qwen2.5-VL 7B route floor | Archived public baseline | qwen2_5_vl_7b | 634 | 1.28391 | 0.0347846 | — | — | — |
| Qwen2.5-VL 7B route floor | Matched all-patch control | qwen2_5_vl_7b_floor_grt_all | 634 | 1.57571 | 0.0487774 | 1 | 2.46645 | 3.39562 |
| Qwen2.5-VL 7B route floor | Matched quality baseline | qwen2_5_vl_7b_floor_quality_base | 634 | 1.57571 | 0.0487774 | 1 | 2.4373 | 3.4355 |
| Qwen2.5-VL 7B route floor | GRT candidate | grt_qwen2_5_vl_7b_dual_floor_s080_o055_cap48 | 634 | 1.59148 | 0.0489087 | 0.847719 | 2.48043 | 3.38126 |
Not all metrics improve. Qwen 3B GRT reports mean throughput 1.67264 fps versus 1.74994 for its all-patch control (about 4.42% lower), even though its Open MOS, Token F1 and patch reuse improve. Route31 mean request time is essentially unchanged versus its all-patch control. Throughput is the mean of per-request sampled-frame rates, not total frames divided by total campaign time, and these single historical runs do not establish repeated speedup. Patch ratios measure patch projection only, not end-to-end FLOPs.
Original immutable 32-row historical snapshot remains byte-identical to the archived numerical bundle; it is not the current release-policy view. The hold does not rewrite historical evidence or change Educational scores, ranks or GRT gates. Full dataset access, GPU/judge reproduction and manuscript alignment remain separate release checks.