← Interactive project page

Complete protocol-screened leaderboard

Educational scores retain the 20 August historical snapshot. High-Motion v2 compares CPU-rescored archived baselines with one new GRT run; no baseline inference was repeated. Download CSV · Versioned source hashes and comparison limits · Public reproduction code · Corrected-reference reproduction guide.

29 Educational results, 19 corrected High-Motion preview results, and 12 educational GRT comparison rows (60 CSV records). Educational: 634 QA / 317 videos. High-Motion v2: the fixed 1,000-record preview of 3,243 source records; metric-specific valid-reference coverage is shown separately. All numeric cells retain full precision in their title and downloadable data.

Educational High-FPS Videos

29 published methods. Rank by reported Open MOS, then Token F1; missing Open MOS sorts last and is never filled in. Open MOS judge: Qwen/Qwen3-VL-32B-Instruct. The verified GRT profiles use eight sampled frames, not a measured high-FPS operating point.

29 educational methods; 634 items each
RankModelMethod IDItemsOpen MOS ↑Token F1 ↑CER ↓WER ↓Exact match ↑Patch recompute ↓Reference patch compute ↓Sampling density (fps)Mean throughput (fps) ↑Source
1Gemini 3.1 Pro Preview (gemini-3.1-pro-preview)gemini-3.1-pro-preview6341.60410.1596331.268811.241660gemini
2GRT (Qwen2.5-VL 7B, route floor 0.80 / 0.55, cap 48)grt_qwen2_5_vl_7b_dual_floor_s080_o055_cap486341.591480.04890871.064011.0537700.8477190.8477190.007479172.48043open
3GRT (Qwen2.5-VL 3B, threshold 0.3)grt_qwen2_5_vl_3b_t036341.589910.09958371.139151.0835100.8856590.8856590.007479171.67264open
4Qwen3-VL 8B Instructqwen3_vl_8b6341.298110.03603981.058891.051980open
5Qwen2.5-VL 7B Instructqwen2_5_vl_7b6341.283910.03478461.046241.039750open
6Gemini 3.6 Flash (gemini-3.6-flash)gemini-3.6-flash6341.249210.08146111.101681.067960gemini
7Qwen3-VL 2B Instructqwen3_vl_2b6341.145110.03314481.04551.03650open
8Qwen3-VL 4B Instructqwen3_vl_4b6341.13880.03486061.055931.051190open
9Qwen2.5-VL 3B Instructqwen2_5_vl_3b6341.135650.03433881.043931.033990open
10Qwen2.5-VL 72B Instructqwen2_5_vl_72b6340.9447950.02686281.059341.047890open
11Qwen2-VL 2B Instructqwen2_vl_2b6340.7902210.02009331.036921.03340open
12Qwen2.5-VL 32B Instructqwen2_5_vl_32b6340.50.02502761.054481.048560open
13Qwen3-VL 32B Instructqwen3_vl_32b6340.1750790.01846761.052391.050850open
14GRT (LLaVA-OneVision Qwen2 0.5B, verified Route31)grt_llava_onevision_0_5b_hf_route31_t00016340.1198740.01419641.110931.1027300.8671540.8671540.007479174.86514open
15LLaVA-OneVision Qwen2 0.5Bllava_onevision_0_5b6340.1167190.01408791.037651.037560open
16LLaVA-OneVision 1.5 8B Instructllava_onevision_1_5_8b6340.03497481.072431.056860open
17Qwen2-VL 7B Instructqwen2_vl_7b6340.03050961.053091.047880open
18Phi-4 Multimodal Instructphi4_multimodal6340.02752871.030081.027620open
19InternVL3 1Binternvl3_1b6340.02716221.0621.06060open
20InternVL3 8Binternvl3_8b6340.02503031.070461.066010open
21VideoLLaMA3 7Bvideollama3_7b6340.02493691.051061.049750open
22LongVA 7Blongva_7b6340.02466161.046781.042410open
23InternVL3 2Binternvl3_2b6340.02230911.034081.035740open
24InternVL2.5 8Binternvl2_5_8b6340.01885331.057381.057190open
25InternVL2.5 4Binternvl2_5_4b6340.017921.020511.023620open
26InternVL2.5 1Binternvl2_5_1b6340.01784271.031471.03360open
27LLaVA-OneVision Qwen2 7Bllava_onevision_original6340.01752021.040561.039820open
28VideoLLaMA3 2Bvideollama3_2b6340.01584761.02451.026830open
29InternVL2.5 2Binternvl2_5_2b6340.01192981.017291.020370open

High-Motion v2: right-ring reference correction, preview-1000

19 methods: 18 archived baselines rescored on CPU and one new HF 0.5B GRT run. All use the same fixed first 1,000 source records, not a full 3,243-record evaluation. The reference target is rightRingFingerMetacarpal, the named right-palm/ring-finger-base proxy; questions and original sampled positions are unchanged. This versioned correction does not claim to recover the original annotation constructor.

Rank by Grid Accuracy descending; exact ties share a competition rank and are ordered by method ID. Null scores are unranked. Every metric is a macro mean over its defined per-record values, with visible metric-specific row and slot/edge coverage. A dash means undefined under the reference mask, not zero. Invalid slots never shift later predictions; FDE uses the original final slot, and transitions require adjacent valid original slots. Token F1 uses canonical label bags on valid positions, with surplus outputs penalized. All 1,000 records remain counted even when no positions are scoreable.

New GRT predictions versus rescored, configuration-checked archived predictions; archived baseline weight revisions and consumed-tensor identity are unproven. This is not a freshly rerun byte-identical paired experiment or a statistical-significance claim.

19 corrected-reference preview methods; eight slots per record
RankModelMethod IDCached/new recordsGrid Accuracy ↑Grid Accuracy ↑ coverageGrid ADE ↓Grid ADE ↓ coverageGrid FDE ↓Grid FDE ↓ coverageTransition Accuracy ↑Transition Accuracy ↑ coverageToken F1 ↑Token F1 ↑ coveragePrediction provenance
1LLaVA-OneVision-2-8B-Instructllava_onevision_2_8b10000.512279861 rows / 6015 slots0.279313861 rows / 6015 slots0.206212669 rows / 669 slots0.731701852 rows / 5101 edges0.513393861 rows / 6015 slotsarchived_baseline
2Qwen3-VL-8B-Instructqwen3_vl_8b10000.426745861 rows / 6015 slots0.357624861 rows / 6015 slots0.310443669 rows / 669 slots0.734879852 rows / 5101 edges0.426938861 rows / 6015 slotsarchived_baseline
3LLaVA-OneVision-1.5-8B-Instructllava_onevision_1_5_8b10000.41958861 rows / 6015 slots0.374625861 rows / 6015 slots0.35337669 rows / 669 slots0.729748852 rows / 5101 edges0.41958861 rows / 6015 slotsarchived_baseline
4VideoLLaMA3-7Bvideollama3_7b10000.416147861 rows / 6015 slots0.479898861 rows / 6015 slots0.591313669 rows / 669 slots0.440789852 rows / 5101 edges0.462768861 rows / 6015 slotsarchived_baseline
5Qwen3-VL-32B-Instructqwen3_vl_32b10000.360901861 rows / 6015 slots0.384434861 rows / 6015 slots0.392697669 rows / 669 slots0.72526852 rows / 5101 edges0.363749861 rows / 6015 slotsarchived_baseline
6Qwen3-VL-4B-Instructqwen3_vl_4b10000.359986861 rows / 6015 slots0.359877861 rows / 6015 slots0.368279669 rows / 669 slots0.735113852 rows / 5101 edges0.359986861 rows / 6015 slotsarchived_baseline
7Qwen3-VL-2B-Instructqwen3_vl_2b10000.275784861 rows / 6015 slots0.423396861 rows / 6015 slots0.471795669 rows / 669 slots0.313081852 rows / 5101 edges0.35428861 rows / 6015 slotsarchived_baseline
8Qwen2.5-VL-72B-Instructqwen2_5_vl_72b10000.235602861 rows / 6015 slots0.513407861 rows / 6015 slots0.569024669 rows / 669 slots0.648879852 rows / 5101 edges0.258947861 rows / 6015 slotsarchived_baseline
9Qwen2.5-VL-32B-Instructqwen2_5_vl_32b10000.176484861 rows / 6015 slots0.505219861 rows / 6015 slots0.504983669 rows / 669 slots0.56006852 rows / 5101 edges0.189601861 rows / 6015 slotsarchived_baseline
10LongVA-7Blongva_7b10000.0760605861 rows / 6015 slots0.770476861 rows / 6015 slots0.632581669 rows / 669 slots0.120532852 rows / 5101 edges0.155908861 rows / 6015 slotsarchived_baseline
11Phi-4-multimodal-instructphi4_multimodal10000.0722015861 rows / 6015 slots1.1359861 rows / 6015 slots1.21451669 rows / 669 slots0.0180751852 rows / 5101 edges0.11382861 rows / 6015 slotsarchived_baseline
12Qwen2-VL-2B-Instructqwen2_vl_2b10000.0656546861 rows / 6015 slots0.771788861 rows / 6015 slots0.625801669 rows / 669 slots0.225872852 rows / 5101 edges0.105367861 rows / 6015 slotsarchived_baseline
13LLaVA-OneVision HF 7Bllava_onevision_original10000.0624772861 rows / 6015 slots0.881176861 rows / 6015 slots0.846691669 rows / 669 slots0.185616852 rows / 5101 edges0.116204861 rows / 6015 slotsarchived_baseline
14VideoLLaMA3-2Bvideollama3_2b10000.0549375861 rows / 6015 slots1.12491861 rows / 6015 slots1.15021669 rows / 669 slots0.000503018852 rows / 5101 edges0.115104861 rows / 6015 slotsarchived_baseline
15Qwen2-VL-7B-Instructqwen2_vl_7b10000.0524114861 rows / 6015 slots0.855721861 rows / 6015 slots0.905482669 rows / 669 slots0.30432852 rows / 5101 edges0.069898861 rows / 6015 slotsarchived_baseline
16GRT · LLaVA-OneVision HF 0.5B (motion SSIM 0.001)grt_llava_hf_0_5b_motion_ssim_t000110000.0495036861 rows / 6015 slots1.0862861 rows / 6015 slots1.06721669 rows / 669 slots0.00534038852 rows / 5101 edges0.0450525861 rows / 6015 slotsnew_grt
17LLaVA-OneVision HF 0.5Bllava_onevision_0_5b10000.0468683861 rows / 6015 slots1.09641861 rows / 6015 slots1.07772669 rows / 669 slots0.0056841852 rows / 5101 edges0.0418983861 rows / 6015 slotsarchived_baseline
18Qwen2.5-VL-3B-Instructqwen2_5_vl_3b10000.0150863861 rows / 6015 slots0.932863861 rows / 6015 slots0.998825669 rows / 669 slots0.576009852 rows / 5101 edges0.0201869861 rows / 6015 slotsarchived_baseline
19Qwen2.5-VL-7B-Instructqwen2_5_vl_7b10000.0149204861 rows / 6015 slots1.18876861 rows / 6015 slots1.30077669 rows / 669 slots0.579396852 rows / 5101 edges0.0135712861 rows / 6015 slotsarchived_baseline

GRT versus its corresponding HF 0.5B baseline

GRT exceeds the corresponding HF 0.5B baseline on observed Grid Accuracy. Point tolerance: 1e-12. Every observed difference is retained, including regressions. This is a point-estimate comparison, not a statistical-significance claim.

All five corrected-reference GRT versus baseline differences
MetricRescored HF 0.5B baselineGRTGRT minus baselineOriented improvement (positive is better)Exceeds point tolerance
Grid Accuracy ↑0.04686830.04950360.002635360.00263536yes
Grid ADE ↓1.096411.0862-0.01020740.0102074yes
Grid FDE ↓1.077721.06721-0.0105010.010501yes
Transition Accuracy ↑0.00568410.00534038-0.000343729-0.000343729no
Token F1 ↑0.04189830.04505250.003154140.00315414yes

Legacy High-Motion results remain withheld

The old target/reference hold and immutable 27-run historical protocol audit remain intact. No old High-Motion score is mixed into the corrected v2 table or its CSV cohort. Reference, scorer, prediction and release provenance are available in the versioned numeric audit.

GRT vs archived and matched controls

All three promoted educational candidates exceed every contracted Open MOS and Token F1 floor, with 11.43–15.23% fewer patch projections than the matched all-patch route. LLaVA-OneVision 7B failed its MOS gate and is not promoted. These are observed point estimates, not statistical-significance claims. Archived Qwen baselines differ from stronger matched controls: their entire gap must not be attributed to GRT.

All 12 educational comparison rows; 634 items and eight sampled frames per method
FamilyControlMethod IDItemsOpen MOS ↑Token F1 ↑Patch recompute ↓Mean throughput (fps) ↑Mean request time (s) ↓
LLaVA-OneVision 0.5B Route31Archived public baselinellava_onevision_0_5b6340.1167190.0140879
LLaVA-OneVision 0.5B Route31Matched quality baselinellava_onevision_0_5b_hf_route31_base6340.1167190.013981514.819282.0401
LLaVA-OneVision 0.5B Route31Matched all-patch controlllava_onevision_0_5b_hf_route31_exact6340.1167190.013981514.82392.00908
LLaVA-OneVision 0.5B Route31GRT candidategrt_llava_onevision_0_5b_hf_route31_t00016340.1198740.01419640.8671544.865142.00911
Qwen2.5-VL 3B t03Archived public baselineqwen2_5_vl_3b6341.135650.0343388
Qwen2.5-VL 3B t03Matched all-patch controlqwen2_5_vl_3b_grt_all6341.558360.099135811.749945.42461
Qwen2.5-VL 3B t03Matched quality baselineqwen2_5_vl_3b_quality_base6341.553630.098891611.631725.88014
Qwen2.5-VL 3B t03GRT candidategrt_qwen2_5_vl_3b_t036341.589910.09958370.8856591.672645.70157
Qwen2.5-VL 7B route floorArchived public baselineqwen2_5_vl_7b6341.283910.0347846
Qwen2.5-VL 7B route floorMatched all-patch controlqwen2_5_vl_7b_floor_grt_all6341.575710.048777412.466453.39562
Qwen2.5-VL 7B route floorMatched quality baselineqwen2_5_vl_7b_floor_quality_base6341.575710.048777412.43733.4355
Qwen2.5-VL 7B route floorGRT candidategrt_qwen2_5_vl_7b_dual_floor_s080_o055_cap486341.591480.04890870.8477192.480433.38126

Not all metrics improve. Qwen 3B GRT reports mean throughput 1.67264 fps versus 1.74994 for its all-patch control (about 4.42% lower), even though its Open MOS, Token F1 and patch reuse improve. Route31 mean request time is essentially unchanged versus its all-patch control. Throughput is the mean of per-request sampled-frame rates, not total frames divided by total campaign time, and these single historical runs do not establish repeated speedup. Patch ratios measure patch projection only, not end-to-end FLOPs.

Original immutable 32-row historical snapshot remains byte-identical to the archived numerical bundle; it is not the current release-policy view. The hold does not rewrite historical evidence or change Educational scores, ranks or GRT gates. Full dataset access, GPU/judge reproduction and manuscript alignment remain separate release checks.