{"meta":{"generated_utc":"2026-08-15T15:00:26Z","name":"ΦBENCH","name_full":"frontier llm infra bench","paper_title":"Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?","counts":{"kfc":55,"lh":20,"e2e":10,"total":85,"all_tasks":85},"weights":{"kfc":55,"lh":20,"e2e":10},"best_overall_pct":63.61,"formats":[{"key":"kfc","name":"KFC · Kernel Function Completion","scope":"single file","submit":"single","n":55},{"key":"lh","name":"LHI · Long-Horizon Implementation","scope":"multi-file","submit":"multiple","n":20},{"key":"e2e","name":"E2EO · End-to-End Optimization","scope":"whole repo","submit":"multiple","n":10}],"formula_perf":"performance: r_new = max(0, 2·r_old − 1), r_old = min(1, 0.5·ln(speedup)/ln(ref_speedup)); hard-gate fail or speedup ≤ 1 → 0","formula_impl":"implementation (binary): all test cases pass with no cheating → 1.0, otherwise 0.0","selection_e2e_perf":"e2e performance tasks: take the non-zero minimum (all-zero → 0)","selection_exception":"exception (kimi-k3, max, e2e perf): if 2nd-best/best < 0.7 treat as an anomalous run → take best; otherwise take 2nd-best","weighting_note":"overall = (kfc_mean×55 + lh_mean×20 + e2e_mean×10) / 85, weighted by task count","intro":"A new question: can a large language model engineer the very infrastructure that powers it? Φ-Bench evaluates that ability on long-horizon, open-ended work over real systems — reading a codebase, locating bottlenecks, and iteratively implementing, profiling, and debugging across the LLM training, inference, and systems stack.","institutions":[{"name":"StepFun","logo":"assets/logos/color/stepfun.png","h":74},{"name":"University of Science and Technology of China","logo":"assets/logos/color/ustc.svg"},{"name":"Peking University","logo":"assets/logos/color/peking.svg"},{"name":"The Hong Kong University of Science and Technology","logo":"assets/logos/color/hkust.png","h":54},{"name":"Yale University","logo":"assets/logos/color/yale.svg"},{"name":"University of Pennsylvania","logo":"assets/logos/color/upenn.svg"}],"links":{"github":"https://github.com/one2piece2hello/LLM-Infra-Bench-faibench","arxiv":"","huggingface":"https://huggingface.co/datasets/faibench-Frontier-Infra-Bench/faibench_Frontier_Infra_Bench"}},"models":["opus-5","kimi-k3","gpt-5.6","glm-5.2","sonnet-5","qwen3.7","qwen3.8","deepseek"],"colors":{"opus-5":"#00e5ff","kimi-k3":"#ff2bd6","gpt-5.6":"#a6ff00","glm-5.2":"#ffb300","sonnet-5":"#7c5cff","qwen3.7":"#ff5c7c","qwen3.8":"#ff8a3d","deepseek":"#5b8cff"},"model_meta":{"opus-5":{"name":"Claude Opus 5","harness":"Claude Code"},"kimi-k3":{"name":"Kimi K3","harness":"Claude Code"},"gpt-5.6":{"name":"GPT-5.6 Sol","harness":"Codex"},"glm-5.2":{"name":"GLM-5.2","harness":"Claude Code"},"sonnet-5":{"name":"Claude Sonnet 5","harness":"Claude Code"},"qwen3.7":{"name":"Qwen3.7 Max","harness":"Claude Code"},"qwen3.8":{"name":"Qwen3.8 Max","harness":"Claude Code"},"deepseek":{"name":"DeepSeek V4 Pro","harness":"Claude Code"}},"topics":{"order":["A","B","C","D","E","F","G","H","I"],"en":{"A":"Training & Post-Training","B":"Inference & Serving","C":"Quantization / Sparsity / Compression","D":"Kernel / Compiler / Runtime","E":"Communication / Interconnect / Data Movement","F":"Hardware / Accelerators / Edge","G":"Data / Retrieval / Storage / Formats","H":"Platform / Observability / Eval / Ops","I":"Cross-cutting (Correctness / Fault-tol / Security)"},"counts":{"A":19,"B":11,"C":6,"D":28,"E":7,"F":3,"G":2,"H":5,"I":4}},"topic_table":{"rows":[{"code":"A","name":"Training & Post-Training Systems","kfc":15,"lh":1,"e2e":3,"total":19,"mid":11},{"code":"B","name":"Inference & Serving Systems","kfc":7,"lh":2,"e2e":2,"total":11,"mid":8},{"code":"C","name":"Low-Precision, Sparsity & Compression","kfc":2,"lh":4,"e2e":0,"total":6,"mid":3},{"code":"D","name":"Kernel, Compiler & Runtime","kfc":19,"lh":8,"e2e":1,"total":28,"mid":9},{"code":"E","name":"Communication, Interconnect & Data Movement","kfc":5,"lh":0,"e2e":2,"total":7,"mid":4},{"code":"F","name":"Hardware, Accelerators, Near-Memory & Edge","kfc":2,"lh":1,"e2e":0,"total":3,"mid":3},{"code":"G","name":"Data, Retrieval, Storage & Formats","kfc":1,"lh":0,"e2e":1,"total":2,"mid":2},{"code":"H","name":"Platform, Observability, Eval & Ops","kfc":3,"lh":1,"e2e":1,"total":5,"mid":4},{"code":"I","name":"Cross-cutting: Correctness / Fault-tolerance / Security / Foundations","kfc":1,"lh":3,"e2e":0,"total":4,"mid":4}],"totals":{"kfc":55,"lh":20,"e2e":10,"total":85,"mid":48},"mid_total":62},"leaderboard":{"max":[{"model":"opus-5","name":"Claude Opus 5","harness":"Claude Code","effort":"max","kfc":0.3716,"lh":0.216,"e2e":0.6294,"total":0.3653,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":1},{"model":"kimi-k3","name":"Kimi K3","harness":"Claude Code","effort":"max","kfc":0.2609,"lh":0.1955,"e2e":0.5641,"total":0.2812,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":2},{"model":"qwen3.8","name":"Qwen3.8 Max","harness":"Claude Code","effort":"max","kfc":0.2879,"lh":0.1661,"e2e":0.441,"total":0.2773,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":3},{"model":"gpt-5.6","name":"GPT-5.6 Sol","harness":"Codex","effort":"max","kfc":0.2546,"lh":0.1398,"e2e":0.4033,"total":0.2451,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":4},{"model":"glm-5.2","name":"GLM-5.2","harness":"Claude Code","effort":"max","kfc":0.2435,"lh":0.1365,"e2e":0.2511,"total":0.2192,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":5},{"model":"sonnet-5","name":"Claude Sonnet 5","harness":"Claude Code","effort":"max","kfc":0.1814,"lh":0.1346,"e2e":0.2274,"total":0.1758,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":6},{"model":"qwen3.7","name":"Qwen3.7 Max","harness":"Claude Code","effort":"max","kfc":0.1647,"lh":0.1279,"e2e":0.204,"total":0.1607,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":7},{"model":"deepseek","name":"DeepSeek V4 Pro","harness":"Claude Code","effort":"max","kfc":0.1605,"lh":0.1145,"e2e":0.0197,"total":0.1331,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":8}],"high":[{"model":"opus-5","name":"Claude Opus 5","harness":"Claude Code","effort":"high","kfc":0.3579,"lh":0.1501,"e2e":0.6571,"total":0.3442,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":1},{"model":"kimi-k3","name":"Kimi K3","harness":"Claude Code","effort":"high","kfc":0.2448,"lh":0.206,"e2e":0.4687,"total":0.262,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":2},{"model":"glm-5.2","name":"GLM-5.2","harness":"Claude Code","effort":"high","kfc":0.2367,"lh":0.1423,"e2e":0.392,"total":0.2328,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":3},{"model":"qwen3.7","name":"Qwen3.7 Max","harness":"Claude Code","effort":"high","kfc":0.1895,"lh":0.1237,"e2e":0.1653,"total":0.1712,"counts":{"kfc":55,"lh":20,"e2e":9},"rank":4},{"model":"gpt-5.6","name":"GPT-5.6 Sol","harness":"Codex","effort":"high","kfc":0.1569,"lh":0.134,"e2e":0.28,"total":0.166,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":5},{"model":"sonnet-5","name":"Claude Sonnet 5","harness":"Claude Code","effort":"high","kfc":0.112,"lh":0.1332,"e2e":0.2739,"total":0.1361,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":6},{"model":"deepseek","name":"DeepSeek V4 Pro","harness":"Claude Code","effort":"high","kfc":0.1325,"lh":0.0532,"e2e":0.1813,"total":0.1196,"counts":{"kfc":55,"lh":20,"e2e":10},"rank":7}]},"topic_matrix":{"max":{"opus-5":{"A":0.5635,"B":0.2624,"C":0.1575,"D":0.3004,"E":0.5714,"F":0.0387,"G":0.4936,"H":0.3417,"I":0.3221},"kimi-k3":{"A":0.3554,"B":0.3046,"C":0.1321,"D":0.2224,"E":0.3574,"F":0.0362,"G":0.2741,"H":0.4115,"I":0.3907},"gpt-5.6":{"A":0.3992,"B":0.2177,"C":0.0299,"D":0.1773,"E":0.3848,"F":0.0,"G":0.184,"H":0.2275,"I":0.3777},"glm-5.2":{"A":0.2873,"B":0.1916,"C":0.0903,"D":0.1855,"E":0.1914,"F":0.0141,"G":0.2741,"H":0.2207,"I":0.5742},"sonnet-5":{"A":0.3402,"B":0.0723,"C":0.0657,"D":0.1413,"E":0.2058,"F":0.0,"G":0.2353,"H":0.0151,"I":0.3369},"qwen3.7":{"A":0.2737,"B":0.1265,"C":0.0588,"D":0.1323,"E":0.1743,"F":0.0544,"G":0.0,"H":0.0683,"I":0.322},"qwen3.8":{"A":0.4392,"B":0.2986,"C":0.0825,"D":0.2298,"E":0.183,"F":0.0014,"G":0.2654,"H":0.2685,"I":0.4628},"deepseek":{"A":0.2781,"B":0.0391,"C":0.0053,"D":0.1309,"E":0.037,"F":0.0086,"G":0.0,"H":0.0703,"I":0.3168}},"high":{"opus-5":{"A":0.5136,"B":0.387,"C":0.1444,"D":0.2645,"E":0.4286,"F":0.0362,"G":0.4286,"H":0.2334,"I":0.4595},"kimi-k3":{"A":0.4393,"B":0.106,"C":0.0835,"D":0.2181,"E":0.4059,"F":0.027,"G":0.3184,"H":0.2216,"I":0.3709},"gpt-5.6":{"A":0.2996,"B":0.0912,"C":0.0297,"D":0.1125,"E":0.3461,"F":0.0,"G":0.087,"H":0.0028,"I":0.3691},"glm-5.2":{"A":0.3896,"B":0.2329,"C":0.0625,"D":0.1727,"E":0.3334,"F":0.0151,"G":0.219,"H":0.0807,"I":0.3477},"sonnet-5":{"A":0.244,"B":0.0956,"C":0.0524,"D":0.1079,"E":0.1609,"F":0.0006,"G":0.0626,"H":0.0059,"I":0.315},"qwen3.7":{"A":0.279,"B":0.038,"C":0.014,"D":0.1413,"E":0.2252,"F":0.0055,"G":0.2097,"H":0.295,"I":0.3372},"qwen3.8":{"A":null,"B":null,"C":null,"D":null,"E":null,"F":null,"G":null,"H":null,"I":null},"deepseek":{"A":0.1746,"B":0.0399,"C":0.0081,"D":0.1384,"E":0.1801,"F":0.0,"G":0.0,"H":0.1914,"I":0.0668}}},"effort":{"tiers":["low","medium","high","xhigh","max"],"models":["opus-5","kimi-k3","gpt-5.6"],"data":{"opus-5":[0.3133,0.3143,0.3191,0.3469,0.3538],"kimi-k3":[0.1745,null,0.2935,null,0.3183],"gpt-5.6":[0.1602,0.2218,0.1827,0.2159,0.2276]},"best":{"opus-5":"max","kimi-k3":"max","gpt-5.6":"max"},"spread":{"opus-5":0.0405,"kimi-k3":0.1438,"gpt-5.6":0.0674},"n_tasks":30},"iterations":{"e2e":[{"task_id":"e2e-a3-moe-train-budget","short_name":"a3-moe-train-budget","infra_name":"moe_train_wallclock_budget","summary":"You are given the complete karpathy/nanoGPT training system (/app/repo, freely modifiable) and a pre-tokenized WikiText-103 corpus, on a single H20. Within the fixed wall-clock budget set by the framework, train the best MoE language model you can, while satisfying a hard lower bound on total parameter count (parameters are re-counted at scoring time). The input is the modifiable training system and the corpus; the output is the trained checkpoint, and scoring looks only at val_bpb on the held-out set (lower is better). Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 1.53375): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress metric is `val_bpb` (held-out bits-per-byte, lower is better); the scored `speedup` is baseline_bpb / candidate `val_bpb` at the fixed wall-clock budget.","bench":"e2e","big_topic":"A","medium_topic":"A3","task_class":"performance","is_performance":true,"metric":"val_bpb","lower_better":true,"models":{"opus-5":{"oracle":null,"traj":[{"n":1,"raw":1.6324,"best":1.6324,"ok":true},{"n":2,"raw":1.4287,"best":1.4287,"ok":true},{"n":3,"raw":1.2279,"best":1.2279,"ok":true},{"n":4,"raw":1.2232,"best":1.2232,"ok":true},{"n":5,"raw":1.2134,"best":1.2134,"ok":true},{"n":6,"raw":1.2054,"best":1.2054,"ok":true},{"n":7,"raw":1.1999,"best":1.1999,"ok":true}],"n_max":16},"kimi-k3":{"oracle":null,"traj":[{"n":1,"raw":2.0766,"best":2.0766,"ok":true},{"n":2,"raw":2.2271,"best":2.0766,"ok":true},{"n":3,"raw":2.1342,"best":2.0766,"ok":true},{"n":4,"raw":2.173,"best":2.0766,"ok":true},{"n":5,"raw":1.4694,"best":1.4694,"ok":true},{"n":6,"raw":1.4713,"best":1.4694,"ok":true},{"n":7,"raw":1.4675,"best":1.4675,"ok":true},{"n":8,"raw":1.4633,"best":1.4633,"ok":true},{"n":9,"raw":1.4469,"best":1.4469,"ok":true},{"n":10,"raw":1.4408,"best":1.4408,"ok":true},{"n":11,"raw":1.4504,"best":1.4408,"ok":true},{"n":12,"raw":1.3065,"best":1.3065,"ok":true},{"n":13,"raw":1.2797,"best":1.2797,"ok":true}],"n_max":16},"gpt-5.6":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":true},{"n":2,"raw":null,"best":null,"ok":true},{"n":3,"raw":null,"best":null,"ok":true},{"n":4,"raw":null,"best":null,"ok":true},{"n":5,"raw":null,"best":null,"ok":true},{"n":6,"raw":null,"best":null,"ok":true},{"n":7,"raw":null,"best":null,"ok":true},{"n":8,"raw":null,"best":null,"ok":true},{"n":9,"raw":null,"best":null,"ok":true},{"n":10,"raw":null,"best":null,"ok":true},{"n":11,"raw":null,"best":null,"ok":true},{"n":12,"raw":null,"best":null,"ok":true},{"n":13,"raw":null,"best":null,"ok":true},{"n":14,"raw":null,"best":null,"ok":true},{"n":15,"raw":null,"best":null,"ok":true},{"n":16,"raw":null,"best":null,"ok":true}],"n_max":null},"glm-5.2":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":false},{"n":2,"raw":1.6444,"best":1.6444,"ok":true},{"n":3,"raw":1.6278,"best":1.6278,"ok":true},{"n":4,"raw":1.6188,"best":1.6188,"ok":true},{"n":5,"raw":1.6112,"best":1.6112,"ok":true},{"n":6,"raw":1.6204,"best":1.6112,"ok":true}],"n_max":16},"sonnet-5":{"oracle":null,"traj":[{"n":1,"raw":2.2423,"best":2.2423,"ok":true},{"n":2,"raw":2.2423,"best":2.2423,"ok":false},{"n":3,"raw":2.1697,"best":2.1697,"ok":true},{"n":4,"raw":2.123,"best":2.123,"ok":true},{"n":5,"raw":2.0691,"best":2.0691,"ok":true},{"n":6,"raw":2.0228,"best":2.0228,"ok":true},{"n":7,"raw":1.9874,"best":1.9874,"ok":true},{"n":8,"raw":1.898,"best":1.898,"ok":true},{"n":9,"raw":1.9254,"best":1.898,"ok":true}],"n_max":16},"qwen3.7":{"oracle":null,"traj":[{"n":1,"raw":1.8647,"best":1.8647,"ok":true},{"n":2,"raw":1.9351,"best":1.8647,"ok":true},{"n":3,"raw":1.815,"best":1.815,"ok":true},{"n":4,"raw":1.7931,"best":1.7931,"ok":true},{"n":5,"raw":1.7931,"best":1.7931,"ok":false},{"n":6,"raw":1.8046,"best":1.7931,"ok":true},{"n":7,"raw":1.812,"best":1.7931,"ok":true},{"n":8,"raw":1.8925,"best":1.7931,"ok":true},{"n":9,"raw":1.8167,"best":1.7931,"ok":true},{"n":10,"raw":1.9107,"best":1.7931,"ok":true},{"n":11,"raw":1.8769,"best":1.7931,"ok":true},{"n":12,"raw":1.7098,"best":1.7098,"ok":true},{"n":13,"raw":1.6562,"best":1.6562,"ok":true},{"n":14,"raw":1.6625,"best":1.6562,"ok":true},{"n":15,"raw":1.7628,"best":1.6562,"ok":true},{"n":16,"raw":1.758,"best":1.6562,"ok":true}],"n_max":16},"qwen3.8":{"oracle":null,"traj":[{"n":1,"raw":2.0362,"best":2.0362,"ok":true},{"n":2,"raw":2.0286,"best":2.0286,"ok":true},{"n":3,"raw":1.9466,"best":1.9466,"ok":true},{"n":4,"raw":1.9967,"best":1.9466,"ok":true},{"n":5,"raw":1.937,"best":1.937,"ok":true},{"n":6,"raw":1.8776,"best":1.8776,"ok":true},{"n":7,"raw":1.825,"best":1.825,"ok":true},{"n":8,"raw":1.8135,"best":1.8135,"ok":true},{"n":9,"raw":1.858,"best":1.8135,"ok":true},{"n":10,"raw":1.7876,"best":1.7876,"ok":true},{"n":11,"raw":1.7334,"best":1.7334,"ok":true},{"n":12,"raw":1.6856,"best":1.6856,"ok":true},{"n":13,"raw":1.7194,"best":1.6856,"ok":true},{"n":14,"raw":1.7322,"best":1.6856,"ok":true},{"n":15,"raw":1.7423,"best":1.6856,"ok":true},{"n":16,"raw":1.6977,"best":1.6856,"ok":true}],"n_max":16},"deepseek":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":false},{"n":2,"raw":null,"best":null,"ok":false},{"n":3,"raw":null,"best":null,"ok":false},{"n":4,"raw":1.874,"best":1.874,"ok":true},{"n":5,"raw":1.8804,"best":1.874,"ok":true},{"n":6,"raw":1.9499,"best":1.874,"ok":true},{"n":7,"raw":1.874,"best":1.874,"ok":false},{"n":8,"raw":1.874,"best":1.874,"ok":false},{"n":9,"raw":1.9549,"best":1.874,"ok":true},{"n":10,"raw":1.9581,"best":1.874,"ok":true},{"n":11,"raw":1.6562,"best":1.6562,"ok":true},{"n":12,"raw":1.68,"best":1.6562,"ok":true},{"n":13,"raw":1.6368,"best":1.6368,"ok":true},{"n":14,"raw":1.8134,"best":1.6368,"ok":true},{"n":15,"raw":1.6366,"best":1.6366,"ok":true},{"n":16,"raw":1.8264,"best":1.6366,"ok":true}],"n_max":16}}},{"task_id":"e2e-a4-token-efficiency-budget","short_name":"a4-token-efficiency-budget","infra_name":"lm_train_token_budget","summary":"Again the complete nanoGPT training system plus the WikiText-103 corpus on a single H20, but the budget is now switched to a fixed number of training tokens. Train the best language model you can without exceeding the given token budget — because the budget is tokens rather than time, merely raising throughput does nothing; you must improve sample efficiency. The input is the modifiable training system and the corpus; the output is a checkpoint, and scoring measures its val_bpb on an unseen held-out set. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 1.20453): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress metric is `dev_val_bpb` (held-out bits-per-byte, lower is better); the scored `speedup` is baseline_bpb / candidate `dev_val_bpb` at the fixed token budget.","bench":"e2e","big_topic":"A","medium_topic":"A4","task_class":"performance","is_performance":true,"metric":"dev_val_bpb","lower_better":true,"models":{"opus-5":{"oracle":null,"traj":[{"n":1,"raw":1.8464,"best":1.8464,"ok":true},{"n":2,"raw":1.7144,"best":1.7144,"ok":true},{"n":3,"raw":1.4708,"best":1.4708,"ok":true},{"n":4,"raw":1.4698,"best":1.4698,"ok":true},{"n":5,"raw":1.4425,"best":1.4425,"ok":true},{"n":6,"raw":1.4092,"best":1.4092,"ok":true},{"n":7,"raw":1.4189,"best":1.4092,"ok":true},{"n":8,"raw":1.4082,"best":1.4082,"ok":true}],"n_max":16},"kimi-k3":{"oracle":null,"traj":[{"n":1,"raw":2.1933,"best":2.1933,"ok":true},{"n":2,"raw":1.9818,"best":1.9818,"ok":true},{"n":3,"raw":1.8543,"best":1.8543,"ok":true},{"n":4,"raw":1.8678,"best":1.8543,"ok":true},{"n":5,"raw":1.7826,"best":1.7826,"ok":true},{"n":6,"raw":1.7829,"best":1.7826,"ok":true},{"n":7,"raw":1.5642,"best":1.5642,"ok":true},{"n":8,"raw":1.5152,"best":1.5152,"ok":true},{"n":9,"raw":1.4238,"best":1.4238,"ok":true},{"n":10,"raw":1.4208,"best":1.4208,"ok":true},{"n":11,"raw":1.4149,"best":1.4149,"ok":true}],"n_max":16},"gpt-5.6":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":true},{"n":2,"raw":null,"best":null,"ok":true},{"n":3,"raw":null,"best":null,"ok":true},{"n":4,"raw":null,"best":null,"ok":true},{"n":5,"raw":null,"best":null,"ok":true},{"n":6,"raw":null,"best":null,"ok":true},{"n":7,"raw":null,"best":null,"ok":true}],"n_max":null},"glm-5.2":{"oracle":null,"traj":[{"n":1,"raw":1.8959,"best":1.8959,"ok":true},{"n":2,"raw":1.8301,"best":1.8301,"ok":true},{"n":3,"raw":1.8409,"best":1.8301,"ok":true},{"n":4,"raw":1.8329,"best":1.8301,"ok":true},{"n":5,"raw":1.8095,"best":1.8095,"ok":true}],"n_max":16},"sonnet-5":{"oracle":null,"traj":[{"n":1,"raw":2.0859,"best":2.0859,"ok":true},{"n":2,"raw":1.9555,"best":1.9555,"ok":true},{"n":3,"raw":1.9751,"best":1.9555,"ok":true},{"n":4,"raw":1.876,"best":1.876,"ok":true},{"n":5,"raw":1.9035,"best":1.876,"ok":true},{"n":6,"raw":1.8747,"best":1.8747,"ok":true},{"n":7,"raw":1.9201,"best":1.8747,"ok":true},{"n":8,"raw":1.8731,"best":1.8731,"ok":true},{"n":9,"raw":1.8423,"best":1.8423,"ok":true},{"n":10,"raw":1.8393,"best":1.8393,"ok":true},{"n":11,"raw":1.6971,"best":1.6971,"ok":true},{"n":12,"raw":1.7003,"best":1.6971,"ok":true},{"n":13,"raw":1.6488,"best":1.6488,"ok":true},{"n":14,"raw":1.5622,"best":1.5622,"ok":true},{"n":15,"raw":1.4982,"best":1.4982,"ok":true},{"n":16,"raw":1.4733,"best":1.4733,"ok":true}],"n_max":16},"qwen3.7":{"oracle":null,"traj":[{"n":1,"raw":2.1169,"best":2.1169,"ok":true},{"n":2,"raw":1.939,"best":1.939,"ok":true},{"n":3,"raw":2.0331,"best":1.939,"ok":true},{"n":4,"raw":1.9407,"best":1.939,"ok":true},{"n":5,"raw":1.9396,"best":1.939,"ok":true},{"n":6,"raw":1.794,"best":1.794,"ok":true},{"n":7,"raw":1.8808,"best":1.794,"ok":true},{"n":8,"raw":1.8945,"best":1.794,"ok":true},{"n":9,"raw":1.7964,"best":1.794,"ok":true},{"n":10,"raw":1.776,"best":1.776,"ok":true},{"n":11,"raw":null,"best":1.776,"ok":false},{"n":12,"raw":1.7328,"best":1.7328,"ok":true},{"n":13,"raw":1.7111,"best":1.7111,"ok":true},{"n":14,"raw":1.7343,"best":1.7111,"ok":true},{"n":15,"raw":1.709,"best":1.709,"ok":true},{"n":16,"raw":1.6942,"best":1.6942,"ok":true}],"n_max":16},"qwen3.8":{"oracle":null,"traj":[{"n":1,"raw":1.9871,"best":1.9871,"ok":true},{"n":2,"raw":1.8956,"best":1.8956,"ok":true},{"n":3,"raw":1.8726,"best":1.8726,"ok":true},{"n":4,"raw":1.8429,"best":1.8429,"ok":true},{"n":5,"raw":1.8419,"best":1.8419,"ok":true},{"n":6,"raw":1.7062,"best":1.7062,"ok":true},{"n":7,"raw":1.7079,"best":1.7062,"ok":true},{"n":8,"raw":1.6798,"best":1.6798,"ok":true},{"n":9,"raw":1.6713,"best":1.6713,"ok":true},{"n":10,"raw":1.6719,"best":1.6713,"ok":true},{"n":11,"raw":1.6636,"best":1.6636,"ok":true}],"n_max":16},"deepseek":{"oracle":null,"traj":[{"n":1,"raw":2.0049,"best":2.0049,"ok":true},{"n":2,"raw":2.0129,"best":2.0049,"ok":true},{"n":3,"raw":2.0649,"best":2.0049,"ok":true},{"n":4,"raw":2.0644,"best":2.0049,"ok":true},{"n":5,"raw":2.0641,"best":2.0049,"ok":true},{"n":6,"raw":2.0037,"best":2.0037,"ok":true},{"n":7,"raw":2.0565,"best":2.0037,"ok":true},{"n":8,"raw":2.0118,"best":2.0037,"ok":true},{"n":9,"raw":2.0159,"best":2.0037,"ok":true},{"n":10,"raw":2.0082,"best":2.0037,"ok":true},{"n":11,"raw":2.0,"best":2.0,"ok":true},{"n":12,"raw":2.0024,"best":2.0,"ok":true},{"n":13,"raw":1.9999,"best":1.9999,"ok":true},{"n":14,"raw":1.9991,"best":1.9991,"ok":true},{"n":15,"raw":1.9998,"best":1.9991,"ok":true},{"n":16,"raw":2.0037,"best":1.9991,"ok":true}],"n_max":16}}},{"task_id":"e2e-a8-peft-adapter-byte-golf","short_name":"a8-peft-adapter-byte-golf","infra_name":"peft_adapter_byte_budget","summary":"You are given the complete huggingface/peft source, a frozen Qwen2.5-0.5B-Instruct base model, a single H20, and a corpus of real distributed-training system code. Adapt this frozen base as well as you can to the corpus's domain, but the adaptation artifacts you may deliver (adapter weights + loading code) are hard-capped at 320 KiB total bytes. The inputs are the base model and the corpus; the outputs are the adapter artifacts under /app/submission and a build_adapted_model loading function, and scoring looks at the relative gain in held-out cross-entropy. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 1.23266): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress metric is `val_gain_nats` (the held-out cross-entropy improvement, in nats, over the frozen base); the scored `speedup` is the relative-gain ratio built from it.","bench":"e2e","big_topic":"A","medium_topic":"A8","task_class":"performance","is_performance":true,"metric":"val_gain_nats","lower_better":false,"models":{"opus-5":{"oracle":null,"traj":[{"n":1,"raw":0.092145,"best":0.092145,"ok":true},{"n":2,"raw":0.11886,"best":0.11886,"ok":true},{"n":3,"raw":0.12189,"best":0.12189,"ok":true},{"n":4,"raw":0.12063,"best":0.12189,"ok":true}],"n_max":16},"kimi-k3":{"oracle":null,"traj":[{"n":1,"raw":0.098379,"best":0.098379,"ok":true},{"n":2,"raw":0.099988,"best":0.099988,"ok":true},{"n":3,"raw":0.11056,"best":0.11056,"ok":true},{"n":4,"raw":0.11373,"best":0.11373,"ok":true},{"n":5,"raw":0.11403,"best":0.11403,"ok":true},{"n":6,"raw":0.12152,"best":0.12152,"ok":true},{"n":7,"raw":0.12469,"best":0.12469,"ok":true},{"n":8,"raw":0.12551,"best":0.12551,"ok":true},{"n":9,"raw":0.12566,"best":0.12566,"ok":true}],"n_max":16},"gpt-5.6":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":false},{"n":2,"raw":null,"best":null,"ok":false},{"n":3,"raw":null,"best":null,"ok":false},{"n":4,"raw":null,"best":null,"ok":false},{"n":5,"raw":null,"best":null,"ok":false},{"n":6,"raw":null,"best":null,"ok":false},{"n":7,"raw":null,"best":null,"ok":false},{"n":8,"raw":null,"best":null,"ok":false},{"n":9,"raw":null,"best":null,"ok":false},{"n":10,"raw":null,"best":null,"ok":false},{"n":11,"raw":null,"best":null,"ok":false},{"n":12,"raw":null,"best":null,"ok":false},{"n":13,"raw":null,"best":null,"ok":false},{"n":14,"raw":null,"best":null,"ok":false},{"n":15,"raw":null,"best":null,"ok":false},{"n":16,"raw":null,"best":null,"ok":false}],"n_max":null},"glm-5.2":{"oracle":null,"traj":[{"n":1,"raw":0.066072,"best":0.066072,"ok":true},{"n":2,"raw":0.066944,"best":0.066944,"ok":true},{"n":3,"raw":0.084466,"best":0.084466,"ok":true},{"n":4,"raw":0.084637,"best":0.084637,"ok":true},{"n":5,"raw":0.083331,"best":0.084637,"ok":true},{"n":6,"raw":0.083217,"best":0.084637,"ok":true},{"n":7,"raw":0.088841,"best":0.088841,"ok":true},{"n":8,"raw":0.084675,"best":0.088841,"ok":true},{"n":9,"raw":0.083914,"best":0.088841,"ok":true},{"n":10,"raw":0.088803,"best":0.088841,"ok":true},{"n":11,"raw":0.08811,"best":0.088841,"ok":true},{"n":12,"raw":0.092442,"best":0.092442,"ok":true},{"n":13,"raw":0.096446,"best":0.096446,"ok":true},{"n":14,"raw":0.10215,"best":0.10215,"ok":true},{"n":15,"raw":0.10154,"best":0.10215,"ok":true},{"n":16,"raw":0.1032,"best":0.1032,"ok":true}],"n_max":16},"sonnet-5":{"oracle":null,"traj":[{"n":1,"raw":0.069618,"best":0.069618,"ok":true},{"n":2,"raw":0.068682,"best":0.069618,"ok":true},{"n":3,"raw":0.085858,"best":0.085858,"ok":true},{"n":4,"raw":0.086099,"best":0.086099,"ok":true}],"n_max":16},"qwen3.7":{"oracle":null,"traj":[{"n":1,"raw":0.015774,"best":0.015774,"ok":true},{"n":2,"raw":0.056249,"best":0.056249,"ok":true},{"n":3,"raw":0.061618,"best":0.061618,"ok":true},{"n":4,"raw":0.060341,"best":0.061618,"ok":true},{"n":5,"raw":0.066212,"best":0.066212,"ok":true},{"n":6,"raw":0.050539,"best":0.066212,"ok":true},{"n":7,"raw":0.061805,"best":0.066212,"ok":true},{"n":8,"raw":0.05929,"best":0.066212,"ok":true},{"n":9,"raw":0.067519,"best":0.067519,"ok":true},{"n":10,"raw":0.061855,"best":0.067519,"ok":true},{"n":11,"raw":0.068167,"best":0.068167,"ok":true},{"n":12,"raw":0.067555,"best":0.068167,"ok":true},{"n":13,"raw":0.066777,"best":0.068167,"ok":true},{"n":14,"raw":0.067391,"best":0.068167,"ok":true},{"n":15,"raw":0.067974,"best":0.068167,"ok":true}],"n_max":16},"qwen3.8":{"oracle":null,"traj":[{"n":1,"raw":0.032818,"best":0.032818,"ok":true},{"n":2,"raw":0.089157,"best":0.089157,"ok":true},{"n":3,"raw":0.0905,"best":0.0905,"ok":true},{"n":4,"raw":0.093516,"best":0.093516,"ok":true},{"n":5,"raw":0.094964,"best":0.094964,"ok":true},{"n":6,"raw":0.095682,"best":0.095682,"ok":true},{"n":7,"raw":0.096015,"best":0.096015,"ok":true},{"n":8,"raw":0.095643,"best":0.096015,"ok":true}],"n_max":16},"deepseek":{"oracle":null,"traj":[{"n":1,"raw":0.042827,"best":0.042827,"ok":true},{"n":2,"raw":0.057069,"best":0.057069,"ok":true},{"n":3,"raw":0.057182,"best":0.057182,"ok":true},{"n":4,"raw":0.048693,"best":0.057182,"ok":true},{"n":5,"raw":0.057463,"best":0.057463,"ok":true},{"n":6,"raw":0.055502,"best":0.057463,"ok":true},{"n":7,"raw":0.056867,"best":0.057463,"ok":true},{"n":8,"raw":0.057963,"best":0.057963,"ok":true},{"n":9,"raw":0.065377,"best":0.065377,"ok":true},{"n":10,"raw":0.069379,"best":0.069379,"ok":true},{"n":11,"raw":0.073476,"best":0.073476,"ok":true},{"n":12,"raw":0.073048,"best":0.073476,"ok":true},{"n":13,"raw":0.073829,"best":0.073829,"ok":true},{"n":14,"raw":0.074416,"best":0.074416,"ok":true},{"n":15,"raw":0.074429,"best":0.074429,"ok":true}],"n_max":16}}},{"task_id":"e2e-b1-kv-traffic-sol","short_name":"kv-traffic-sol","infra_name":"paged_kv_cache_traffic","summary":"You are given a complete vLLM 0.10.1.1 source tree (importable, with torch/triton present) on a single H20 — the task is the movement bandwidth of the paged KV cache. A large fraction of an inference engine's memory bandwidth is spent moving the KV cache, and you must bring these movements as close as possible to the hardware's peak bandwidth. The inputs are the page pool and the per-request block table; the output is the moved KV cache, and scoring also reports what percentage of measured HBM peak each case reaches. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 2.57993): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress metric is `dev_public_score` (a public proxy formed from the geomean SOL-fraction / achieved-HBM-bandwidth over the timed cases); the scored `speedup` is the ABBA-paired ratio of that metric against the strong baseline.","bench":"e2e","big_topic":"B","medium_topic":"B1","task_class":"performance","is_performance":true,"metric":"dev_public_score","lower_better":false,"models":{"opus-5":{"oracle":null,"traj":[{"n":1,"raw":2118.0,"best":2118.0,"ok":true},{"n":2,"raw":2271.0,"best":2271.0,"ok":true},{"n":3,"raw":2265.0,"best":2271.0,"ok":true},{"n":4,"raw":2242.0,"best":2271.0,"ok":true},{"n":5,"raw":2290.0,"best":2290.0,"ok":true},{"n":6,"raw":2274.0,"best":2290.0,"ok":true},{"n":7,"raw":2259.0,"best":2290.0,"ok":true}],"n_max":16},"kimi-k3":{"oracle":null,"traj":[{"n":1,"raw":1970.0,"best":1970.0,"ok":true},{"n":2,"raw":2064.0,"best":2064.0,"ok":true},{"n":3,"raw":2042.0,"best":2064.0,"ok":true},{"n":4,"raw":2055.0,"best":2064.0,"ok":true},{"n":5,"raw":2144.0,"best":2144.0,"ok":true},{"n":6,"raw":2335.0,"best":2335.0,"ok":true},{"n":7,"raw":2337.0,"best":2337.0,"ok":true},{"n":8,"raw":2272.0,"best":2337.0,"ok":true},{"n":9,"raw":2323.0,"best":2337.0,"ok":true}],"n_max":16},"gpt-5.6":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":true},{"n":2,"raw":null,"best":null,"ok":true},{"n":3,"raw":null,"best":null,"ok":true},{"n":4,"raw":null,"best":null,"ok":true},{"n":5,"raw":null,"best":null,"ok":true},{"n":6,"raw":null,"best":null,"ok":true},{"n":7,"raw":null,"best":null,"ok":true}],"n_max":null},"glm-5.2":{"oracle":null,"traj":[{"n":1,"raw":724.2,"best":724.2,"ok":true},{"n":2,"raw":813.3,"best":813.3,"ok":true},{"n":3,"raw":924.6,"best":924.6,"ok":true},{"n":4,"raw":924.1,"best":924.6,"ok":true},{"n":5,"raw":996.2,"best":996.2,"ok":true},{"n":6,"raw":1306.0,"best":1306.0,"ok":true},{"n":7,"raw":2059.0,"best":2059.0,"ok":true},{"n":8,"raw":2109.0,"best":2109.0,"ok":true},{"n":9,"raw":2091.0,"best":2109.0,"ok":true}],"n_max":16},"sonnet-5":{"oracle":null,"traj":[{"n":1,"raw":934.7,"best":934.7,"ok":true},{"n":2,"raw":909.2,"best":934.7,"ok":true},{"n":3,"raw":891.9,"best":934.7,"ok":true}],"n_max":16},"qwen3.7":{"oracle":null,"traj":[{"n":1,"raw":800.2,"best":800.2,"ok":true},{"n":2,"raw":794.1,"best":800.2,"ok":true},{"n":3,"raw":826.9,"best":826.9,"ok":true},{"n":4,"raw":845.9,"best":845.9,"ok":true},{"n":5,"raw":846.5,"best":846.5,"ok":true}],"n_max":16},"qwen3.8":{"oracle":null,"traj":[{"n":1,"raw":1937.0,"best":1937.0,"ok":true},{"n":2,"raw":1990.0,"best":1990.0,"ok":true},{"n":3,"raw":1977.0,"best":1990.0,"ok":true},{"n":4,"raw":1969.0,"best":1990.0,"ok":true},{"n":5,"raw":1966.0,"best":1990.0,"ok":true},{"n":6,"raw":1982.0,"best":1990.0,"ok":true},{"n":7,"raw":2005.0,"best":2005.0,"ok":true},{"n":8,"raw":1952.0,"best":2005.0,"ok":true},{"n":9,"raw":2001.0,"best":2005.0,"ok":true},{"n":10,"raw":1971.0,"best":2005.0,"ok":true},{"n":11,"raw":1774.0,"best":2005.0,"ok":true},{"n":12,"raw":1953.0,"best":2005.0,"ok":true},{"n":13,"raw":2022.0,"best":2022.0,"ok":true}],"n_max":16},"deepseek":{"oracle":null,"traj":[{"n":1,"raw":51.36,"best":51.36,"ok":true},{"n":2,"raw":68.31,"best":68.31,"ok":true},{"n":3,"raw":83.92,"best":83.92,"ok":true},{"n":4,"raw":276.7,"best":276.7,"ok":true},{"n":5,"raw":316.9,"best":316.9,"ok":true},{"n":6,"raw":397.2,"best":397.2,"ok":true},{"n":7,"raw":457.8,"best":457.8,"ok":true},{"n":8,"raw":481.9,"best":481.9,"ok":true},{"n":9,"raw":516.7,"best":516.7,"ok":true},{"n":10,"raw":516.2,"best":516.7,"ok":true},{"n":11,"raw":510.9,"best":516.7,"ok":true}],"n_max":16}}},{"task_id":"e2e-vllm-scheduler-mixed-batch-serving","short_name":"vllm-scheduler-mixed-batch-serving","infra_name":"vllm_continuous_batch_serving","summary":"You are given a vLLM OpenAI-compatible serving instance running on a single H20, up against a strong baseline that has already been hardened on the scheduling side (CUDA graph enabled, chunked-prefill token budget tuned, etc.). Without changing the greedy outputs at all, you must make this service serve the hidden mixed request shapes faster (through request scheduling and continuous batching). The input is the service launched by /app/submission/launch_server.sh; the output is the service's responses, and scoring is the median of baseline_time/your_time over the hidden workload. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 1.28571): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress metric is `dev_tokens_per_s` (serving throughput in tokens/s); the scored `speedup` is the median of baseline_time / your_time over the hidden workloads.","bench":"e2e","big_topic":"B","medium_topic":"B2","task_class":"performance","is_performance":true,"metric":"dev_tokens_per_s","lower_better":false,"models":{"opus-5":{"oracle":null,"traj":[{"n":1,"raw":4084.2,"best":4084.2,"ok":true},{"n":2,"raw":4057.6,"best":4084.2,"ok":true},{"n":3,"raw":3643.8,"best":4084.2,"ok":true},{"n":4,"raw":3739.4,"best":4084.2,"ok":true},{"n":5,"raw":4106.0,"best":4106.0,"ok":true}],"n_max":16},"kimi-k3":{"oracle":null,"traj":[{"n":1,"raw":3667.1,"best":3667.1,"ok":true},{"n":2,"raw":3886.7,"best":3886.7,"ok":true},{"n":3,"raw":3937.2,"best":3937.2,"ok":true},{"n":4,"raw":3793.3,"best":3937.2,"ok":true},{"n":5,"raw":3851.0,"best":3937.2,"ok":true},{"n":6,"raw":3988.2,"best":3988.2,"ok":true},{"n":7,"raw":4091.4,"best":4091.4,"ok":true},{"n":8,"raw":4300.3,"best":4300.3,"ok":true},{"n":9,"raw":3797.9,"best":4300.3,"ok":true},{"n":10,"raw":3999.2,"best":4300.3,"ok":true},{"n":11,"raw":3829.5,"best":4300.3,"ok":true},{"n":12,"raw":4096.6,"best":4300.3,"ok":true},{"n":13,"raw":4097.9,"best":4300.3,"ok":true},{"n":14,"raw":3903.6,"best":4300.3,"ok":true},{"n":15,"raw":3937.4,"best":4300.3,"ok":true},{"n":16,"raw":4007.3,"best":4300.3,"ok":true}],"n_max":16},"gpt-5.6":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":false},{"n":2,"raw":null,"best":null,"ok":false},{"n":3,"raw":null,"best":null,"ok":false},{"n":4,"raw":null,"best":null,"ok":false},{"n":5,"raw":null,"best":null,"ok":false}],"n_max":null},"glm-5.2":{"oracle":null,"traj":[{"n":1,"raw":3396.4,"best":3396.4,"ok":true},{"n":2,"raw":3275.9,"best":3396.4,"ok":true},{"n":3,"raw":3504.1,"best":3504.1,"ok":true},{"n":4,"raw":3902.7,"best":3902.7,"ok":true},{"n":5,"raw":3902.3,"best":3902.7,"ok":true},{"n":6,"raw":3898.1,"best":3902.7,"ok":true},{"n":7,"raw":3946.6,"best":3946.6,"ok":true}],"n_max":16},"sonnet-5":{"oracle":null,"traj":[{"n":1,"raw":3399.6,"best":3399.6,"ok":true},{"n":2,"raw":3480.2,"best":3480.2,"ok":true},{"n":3,"raw":3577.7,"best":3577.7,"ok":true},{"n":4,"raw":3651.8,"best":3651.8,"ok":true},{"n":5,"raw":3511.5,"best":3651.8,"ok":true},{"n":6,"raw":0.0,"best":3651.8,"ok":false},{"n":7,"raw":3422.6,"best":3651.8,"ok":true},{"n":8,"raw":3633.1,"best":3651.8,"ok":true},{"n":9,"raw":3582.4,"best":3651.8,"ok":true},{"n":10,"raw":3611.9,"best":3651.8,"ok":true},{"n":11,"raw":3587.3,"best":3651.8,"ok":true},{"n":12,"raw":3522.6,"best":3651.8,"ok":true},{"n":13,"raw":3573.2,"best":3651.8,"ok":true},{"n":14,"raw":3769.3,"best":3769.3,"ok":true},{"n":15,"raw":3530.5,"best":3769.3,"ok":true},{"n":16,"raw":3642.0,"best":3769.3,"ok":true}],"n_max":16},"qwen3.7":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":false},{"n":2,"raw":3384.1,"best":3384.1,"ok":true},{"n":3,"raw":3384.4,"best":3384.4,"ok":true},{"n":4,"raw":3522.0,"best":3522.0,"ok":true},{"n":5,"raw":3431.9,"best":3522.0,"ok":true},{"n":6,"raw":3332.8,"best":3522.0,"ok":true},{"n":7,"raw":3236.1,"best":3522.0,"ok":true},{"n":8,"raw":3253.3,"best":3522.0,"ok":true},{"n":9,"raw":3564.1,"best":3564.1,"ok":true},{"n":10,"raw":3361.0,"best":3564.1,"ok":true},{"n":11,"raw":3296.3,"best":3564.1,"ok":true},{"n":12,"raw":3608.8,"best":3608.8,"ok":true},{"n":13,"raw":3622.1,"best":3622.1,"ok":true},{"n":14,"raw":3308.3,"best":3622.1,"ok":true},{"n":15,"raw":3596.3,"best":3622.1,"ok":true},{"n":16,"raw":3610.9,"best":3622.1,"ok":true}],"n_max":16},"qwen3.8":{"oracle":null,"traj":[{"n":1,"raw":3279.6,"best":3279.6,"ok":true},{"n":2,"raw":3607.7,"best":3607.7,"ok":true}],"n_max":16},"deepseek":{"oracle":null,"traj":[{"n":1,"raw":3265.1,"best":3265.1,"ok":true},{"n":2,"raw":3152.9,"best":3265.1,"ok":true},{"n":3,"raw":0.0,"best":3265.1,"ok":false},{"n":4,"raw":3249.4,"best":3265.1,"ok":true},{"n":5,"raw":3282.1,"best":3282.1,"ok":true},{"n":6,"raw":3450.5,"best":3450.5,"ok":true},{"n":7,"raw":3520.5,"best":3520.5,"ok":true},{"n":8,"raw":3470.4,"best":3520.5,"ok":true},{"n":9,"raw":3526.9,"best":3526.9,"ok":true},{"n":10,"raw":3556.0,"best":3556.0,"ok":true},{"n":11,"raw":3524.3,"best":3556.0,"ok":true},{"n":12,"raw":3626.1,"best":3626.1,"ok":true},{"n":13,"raw":3443.1,"best":3626.1,"ok":true},{"n":14,"raw":3664.6,"best":3664.6,"ok":true},{"n":15,"raw":3581.3,"best":3664.6,"ok":true},{"n":16,"raw":3615.1,"best":3664.6,"ok":true}],"n_max":16}}},{"task_id":"e2e-d1-varlen-prefill-attn-sol","short_name":"varlen-prefill-attn-sol","infra_name":"varlen_causal_prefill_attn","summary":"You are given a complete vLLM 0.10.1.1 source tree on a single H20 — the task is variable-length causal prefill attention in a continuous-batching engine. A batch of prompts is packed into one flat tensor, with a cumulative-length array describing the boundaries; you must bring this one prefill as close as possible to the hardware limit. The inputs are the packed q/k/v and cu_seqlens; the output is the causal-attention result, and scoring reports each case's per-call latency and the TFLOP/s achieved. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 1.5317): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress metric is `dev_public_score` (a public proxy from each case's per-call latency / achieved TFLOP/s); the scored `speedup` is the ABBA-paired ratio against the strong baseline.","bench":"e2e","big_topic":"D","medium_topic":"D1","task_class":"performance","is_performance":true,"metric":"dev_public_score","lower_better":false,"models":{"opus-5":{"oracle":null,"traj":[{"n":1,"raw":99.24,"best":99.24,"ok":true},{"n":2,"raw":113.1,"best":113.1,"ok":true},{"n":3,"raw":114.7,"best":114.7,"ok":true},{"n":4,"raw":115.5,"best":115.5,"ok":true},{"n":5,"raw":118.3,"best":118.3,"ok":true},{"n":6,"raw":118.1,"best":118.3,"ok":true},{"n":7,"raw":118.6,"best":118.6,"ok":true}],"n_max":16},"kimi-k3":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":false},{"n":2,"raw":86.77,"best":86.77,"ok":true},{"n":3,"raw":88.01,"best":88.01,"ok":true},{"n":4,"raw":100.2,"best":100.2,"ok":true},{"n":5,"raw":102.6,"best":102.6,"ok":true},{"n":6,"raw":102.8,"best":102.8,"ok":true},{"n":7,"raw":104.9,"best":104.9,"ok":true},{"n":8,"raw":110.7,"best":110.7,"ok":true},{"n":9,"raw":110.6,"best":110.7,"ok":true},{"n":10,"raw":111.0,"best":111.0,"ok":true},{"n":11,"raw":113.2,"best":113.2,"ok":true},{"n":12,"raw":113.1,"best":113.2,"ok":true},{"n":13,"raw":113.2,"best":113.2,"ok":true},{"n":14,"raw":113.3,"best":113.3,"ok":true}],"n_max":16},"gpt-5.6":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":false},{"n":2,"raw":null,"best":null,"ok":false},{"n":3,"raw":null,"best":null,"ok":false},{"n":4,"raw":null,"best":null,"ok":true},{"n":5,"raw":null,"best":null,"ok":true},{"n":6,"raw":null,"best":null,"ok":true},{"n":7,"raw":null,"best":null,"ok":true}],"n_max":null},"glm-5.2":{"oracle":null,"traj":[{"n":1,"raw":50.49,"best":50.49,"ok":true},{"n":2,"raw":63.73,"best":63.73,"ok":true},{"n":3,"raw":68.19,"best":68.19,"ok":true},{"n":4,"raw":68.09,"best":68.19,"ok":true}],"n_max":16},"sonnet-5":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":false},{"n":2,"raw":null,"best":null,"ok":false},{"n":3,"raw":87.85,"best":87.85,"ok":true},{"n":4,"raw":88.37,"best":88.37,"ok":true}],"n_max":16},"qwen3.7":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":false},{"n":2,"raw":null,"best":null,"ok":false},{"n":3,"raw":null,"best":null,"ok":false},{"n":4,"raw":52.43,"best":52.43,"ok":true},{"n":5,"raw":52.11,"best":52.43,"ok":true},{"n":6,"raw":52.66,"best":52.66,"ok":true},{"n":7,"raw":52.83,"best":52.83,"ok":true}],"n_max":16},"qwen3.8":{"oracle":null,"traj":[{"n":1,"raw":91.2,"best":91.2,"ok":true},{"n":2,"raw":99.61,"best":99.61,"ok":true},{"n":3,"raw":104.9,"best":104.9,"ok":true}],"n_max":16},"deepseek":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":false},{"n":2,"raw":88.22,"best":88.22,"ok":true},{"n":3,"raw":88.22,"best":88.22,"ok":true}],"n_max":16}}},{"task_id":"correctness-e2e-e5-checkpoint-transfer-integrity","short_name":"checkpoint-transfer-integrity","infra_name":"checkpoint_transfer_integrity","summary":"This is a checkpoint tiered-transfer integrity-and-recovery task, semantically aligned with real systems such as NeMo's S3CheckpointIO. You must implement the CheckpointTransfer class: slice the checkpoint binary into chunks, compute a crc32 per chunk and generate a manifest, perform an all-or-nothing durable upload, verify each chunk on download with resumable-transfer support, and reconstruct a single lost chunk using an XOR parity chunk. The inputs are the checkpoint byte blocks and an externally injected chunk store; the outputs are the results of upload/download/reconstruction and the normalized artifact paths. Metric: binary implementation-class score — 1.0 only if all hidden cases and hard gates pass; any single case failure or a triggered forbidden-edit/cheat condition is 0.0, with no partial credit.","bench":"e2e","big_topic":"E","medium_topic":"E5","task_class":"implementation","is_performance":false,"metric":null,"lower_better":false,"models":{}},{"task_id":"correctness-e2e-e5-tiered-storage-io","short_name":"tiered-storage-io","infra_name":"tiered_storage_engine","summary":"This is a tiered-storage engine and cross-tier consistency task, semantically aligned with real systems such as seaweedfs volume-tier, RocketMQ TieredMessageStore, and Kafka RemoteLogManager. You must implement TieredStore: a capacity-limited in-memory hot tier layered over a durable cold tier, supporting size-driven LRU eviction, resumable segmented migration, and integrity verification. The input is read/write/migration operations in arbitrary order and interleaving; the output must, under any access sequence, be consistent with the reference model's cross-tier read semantics. Metric: binary implementation-class score — 1.0 only if all hidden cases and hard gates pass; any single case failure or a triggered forbidden-edit/cheat condition is 0.0, with no partial credit.","bench":"e2e","big_topic":"E","medium_topic":"E5","task_class":"implementation","is_performance":false,"metric":null,"lower_better":false,"models":{}},{"task_id":"e2e-g2-embed-compress-golf","short_name":"embed-compress-golf","infra_name":"embedding_compress_retrieval","summary":"You are given a complete UKPLab/sentence-transformers checkout and a frozen 384-dimensional all-MiniLM-L6-v2, in an environment with 8 CPU cores, no GPU, and no internet. Under a hard budget of at most 64 bytes per vector (384-dim fp32, or even int8, won't fit, so heavy compression is required), make the nDCG@10 on the held-out retrieval set as high as possible, optionally adding a second-stage rerank over a candidate short-list. The inputs are the texts and queries; the outputs are the encoder (and optional reranker) under /app/submission, and scoring is eval-only, loading only your artifacts for evaluation. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 1.42902): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress metric is `ndcg_at_10` (nDCG@10 on the held-out retrieval set at the fixed per-vector byte budget); the scored `speedup` is candidate `ndcg_at_10` / baseline nDCG@10.","bench":"e2e","big_topic":"G","medium_topic":"G2","task_class":"performance","is_performance":true,"metric":"ndcg_at_10","lower_better":false,"models":{"opus-5":{"oracle":null,"traj":[{"n":1,"raw":0.2424,"best":0.2424,"ok":true},{"n":2,"raw":0.24404,"best":0.24404,"ok":true}],"n_max":16},"kimi-k3":{"oracle":null,"traj":[{"n":1,"raw":0.23075,"best":0.23075,"ok":true},{"n":2,"raw":0.24699,"best":0.24699,"ok":true},{"n":3,"raw":0.25662,"best":0.25662,"ok":true}],"n_max":16},"gpt-5.6":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":true},{"n":2,"raw":null,"best":null,"ok":true},{"n":3,"raw":null,"best":null,"ok":true},{"n":4,"raw":null,"best":null,"ok":true},{"n":5,"raw":null,"best":null,"ok":true},{"n":6,"raw":null,"best":null,"ok":true}],"n_max":null},"glm-5.2":{"oracle":null,"traj":[{"n":1,"raw":0.22847,"best":0.22847,"ok":true},{"n":2,"raw":0.22847,"best":0.22847,"ok":true},{"n":3,"raw":0.25266,"best":0.25266,"ok":true},{"n":4,"raw":0.26386,"best":0.26386,"ok":true},{"n":5,"raw":0.26143,"best":0.26386,"ok":true}],"n_max":16},"sonnet-5":{"oracle":null,"traj":[{"n":1,"raw":0.23075,"best":0.23075,"ok":true}],"n_max":16},"qwen3.7":{"oracle":null,"traj":[{"n":1,"raw":0.22847,"best":0.22847,"ok":true},{"n":2,"raw":0.23075,"best":0.23075,"ok":true},{"n":3,"raw":0.23075,"best":0.23075,"ok":true},{"n":4,"raw":0.22358,"best":0.23075,"ok":true},{"n":5,"raw":0.23075,"best":0.23075,"ok":true},{"n":6,"raw":0.23075,"best":0.23075,"ok":true},{"n":7,"raw":0.23277,"best":0.23277,"ok":true},{"n":8,"raw":0.23927,"best":0.23927,"ok":true},{"n":9,"raw":0.2382,"best":0.23927,"ok":true},{"n":10,"raw":0.23887,"best":0.23927,"ok":true},{"n":11,"raw":0.24401,"best":0.24401,"ok":true},{"n":12,"raw":0.23942,"best":0.24401,"ok":true}],"n_max":16},"qwen3.8":{"oracle":null,"traj":[{"n":1,"raw":0.23075,"best":0.23075,"ok":true},{"n":2,"raw":0.2286,"best":0.23075,"ok":true},{"n":3,"raw":0.24018,"best":0.24018,"ok":true},{"n":4,"raw":0.24137,"best":0.24137,"ok":true}],"n_max":16},"deepseek":{"oracle":null,"traj":[{"n":1,"raw":0.22869,"best":0.22869,"ok":true},{"n":2,"raw":0.23075,"best":0.23075,"ok":true},{"n":3,"raw":0.24926,"best":0.24926,"ok":true},{"n":4,"raw":0.23652,"best":0.24926,"ok":true},{"n":5,"raw":0.23886,"best":0.24926,"ok":true},{"n":6,"raw":0.24367,"best":0.24926,"ok":true}],"n_max":16}}},{"task_id":"e2e-h3-eval-harness-throughput-quality","short_name":"eval-scoring-throughput","infra_name":"eval_harness_scoring_throughput","summary":"You are given a complete EleutherAI/lm-evaluation-harness checkout (editable-installable), in an environment with 8 CPU cores and no GPU — what is under evaluation is the pure-CPU scoring/aggregation pipeline that runs after model inference. You must optimize this end-to-end throughput: regex answer extraction, the take_first / majority_vote transforms, argmax over multiple-choice loglikelihoods, and metric computations such as exact_match / contains / prefix_match. The input is a batch of evaluation records that already carry the model outputs; the output is a {id, score} per sample (the results must match the reference). Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 2.24419): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress metric is `dev_public_score` (a public proxy of the pure-CPU scoring-pipeline throughput); the scored `speedup` is baseline_time / your_time, measured ABBA-paired.","bench":"e2e","big_topic":"H","medium_topic":"H3","task_class":"performance","is_performance":true,"metric":"dev_public_score","lower_better":false,"models":{"opus-5":{"oracle":null,"traj":[{"n":1,"raw":331740.0,"best":331740.0,"ok":true},{"n":2,"raw":403900.0,"best":403900.0,"ok":true},{"n":3,"raw":576880.0,"best":576880.0,"ok":true},{"n":4,"raw":614640.0,"best":614640.0,"ok":true},{"n":5,"raw":651810.0,"best":651810.0,"ok":true},{"n":6,"raw":634140.0,"best":651810.0,"ok":true},{"n":7,"raw":666510.0,"best":666510.0,"ok":true}],"n_max":16},"kimi-k3":{"oracle":null,"traj":[{"n":1,"raw":215440.0,"best":215440.0,"ok":true},{"n":2,"raw":406840.0,"best":406840.0,"ok":true},{"n":3,"raw":457940.0,"best":457940.0,"ok":true},{"n":4,"raw":540470.0,"best":540470.0,"ok":true},{"n":5,"raw":705240.0,"best":705240.0,"ok":true},{"n":6,"raw":1340500.0,"best":1340500.0,"ok":true},{"n":7,"raw":1474000.0,"best":1474000.0,"ok":true},{"n":8,"raw":1399400.0,"best":1474000.0,"ok":true},{"n":9,"raw":1441500.0,"best":1474000.0,"ok":true},{"n":10,"raw":1451700.0,"best":1474000.0,"ok":true},{"n":11,"raw":1752000.0,"best":1752000.0,"ok":true},{"n":12,"raw":1744000.0,"best":1752000.0,"ok":true},{"n":13,"raw":1936900.0,"best":1936900.0,"ok":true}],"n_max":16},"gpt-5.6":{"oracle":null,"traj":[{"n":1,"raw":null,"best":null,"ok":true},{"n":2,"raw":null,"best":null,"ok":true},{"n":3,"raw":null,"best":null,"ok":true},{"n":4,"raw":null,"best":null,"ok":true},{"n":5,"raw":null,"best":null,"ok":true},{"n":6,"raw":null,"best":null,"ok":true},{"n":7,"raw":null,"best":null,"ok":true}],"n_max":null},"glm-5.2":{"oracle":null,"traj":[{"n":1,"raw":313180.0,"best":313180.0,"ok":true},{"n":2,"raw":303490.0,"best":313180.0,"ok":true},{"n":3,"raw":335750.0,"best":335750.0,"ok":true},{"n":4,"raw":369870.0,"best":369870.0,"ok":true},{"n":5,"raw":388250.0,"best":388250.0,"ok":true},{"n":6,"raw":380760.0,"best":388250.0,"ok":true},{"n":7,"raw":390670.0,"best":390670.0,"ok":true},{"n":8,"raw":395420.0,"best":395420.0,"ok":true},{"n":9,"raw":408460.0,"best":408460.0,"ok":true}],"n_max":16},"sonnet-5":{"oracle":null,"traj":[{"n":1,"raw":222290.0,"best":222290.0,"ok":true},{"n":2,"raw":250050.0,"best":250050.0,"ok":true},{"n":3,"raw":209530.0,"best":250050.0,"ok":true},{"n":4,"raw":249090.0,"best":250050.0,"ok":true}],"n_max":16},"qwen3.7":{"oracle":null,"traj":[{"n":1,"raw":329220.0,"best":329220.0,"ok":true},{"n":2,"raw":349270.0,"best":349270.0,"ok":true},{"n":3,"raw":367950.0,"best":367950.0,"ok":true},{"n":4,"raw":367990.0,"best":367990.0,"ok":true},{"n":5,"raw":376950.0,"best":376950.0,"ok":true},{"n":6,"raw":379460.0,"best":379460.0,"ok":true},{"n":7,"raw":382410.0,"best":382410.0,"ok":true}],"n_max":16},"qwen3.8":{"oracle":null,"traj":[{"n":1,"raw":227400.0,"best":227400.0,"ok":true},{"n":2,"raw":385130.0,"best":385130.0,"ok":true},{"n":3,"raw":382010.0,"best":385130.0,"ok":true},{"n":4,"raw":912100.0,"best":912100.0,"ok":true},{"n":5,"raw":1095400.0,"best":1095400.0,"ok":true},{"n":6,"raw":1166900.0,"best":1166900.0,"ok":true},{"n":7,"raw":1168500.0,"best":1168500.0,"ok":true}],"n_max":16},"deepseek":{"oracle":null,"traj":[{"n":1,"raw":210820.0,"best":210820.0,"ok":true},{"n":2,"raw":243920.0,"best":243920.0,"ok":true},{"n":3,"raw":306230.0,"best":306230.0,"ok":true},{"n":4,"raw":332130.0,"best":332130.0,"ok":true},{"n":5,"raw":448510.0,"best":448510.0,"ok":true},{"n":6,"raw":474310.0,"best":474310.0,"ok":true},{"n":7,"raw":516300.0,"best":516300.0,"ok":true},{"n":8,"raw":544560.0,"best":544560.0,"ok":true},{"n":9,"raw":541810.0,"best":544560.0,"ok":true}],"n_max":16}}}],"lh":[{"task_id":"wro-vllm-v1-priority-request-queue","short_name":"wro-vllm-v1-priority-request-queue","infra_name":"vllm_priority_request_queue","summary":"This is the priority request queue of the vLLM V1 scheduler, which dequeues in ascending order of (priority, arrival_time, insertion order). It currently uses an unordered list: enqueue is O(1), but every peek/pop linearly scans for the minimum (draining n items is O(n^2)); you must make it fast while keeping the dequeue order exactly identical. The input is a stream of add/pop/remove operations; the output is the exact dequeue-order trace. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 312): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"B","medium_topic":"B10","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":312.0,"traj":[{"n":1,"raw":304.1,"best":304.1,"ok":true,"reward":0.9874},{"n":2,"raw":321.1,"best":321.1,"ok":true,"reward":1.015},{"n":3,"raw":236.6,"best":321.1,"ok":true,"reward":0.8791},{"n":4,"raw":400.3,"best":400.3,"ok":true,"reward":1.141},{"n":5,"raw":408.3,"best":408.3,"ok":true,"reward":1.154},{"n":6,"raw":482.9,"best":482.9,"ok":true,"reward":1.274},{"n":7,"raw":481.5,"best":482.9,"ok":true,"reward":1.272},{"n":8,"raw":480.7,"best":482.9,"ok":true,"reward":1.27},{"n":9,"raw":510.8,"best":510.8,"ok":true,"reward":1.319}],"final_speedup":498.498181,"vs_oracle":1.597751},"kimi-k3":{"oracle":312.0,"traj":[{"n":1,"raw":100.6,"best":100.6,"ok":true,"reward":0.6613},{"n":2,"raw":253.8,"best":253.8,"ok":true,"reward":0.9067},{"n":3,"raw":253.3,"best":253.8,"ok":true,"reward":0.906},{"n":4,"raw":242.2,"best":253.8,"ok":true,"reward":0.8881},{"n":5,"raw":307.7,"best":307.7,"ok":true,"reward":0.9931},{"n":6,"raw":321.6,"best":321.6,"ok":true,"reward":1.015},{"n":7,"raw":315.2,"best":321.6,"ok":true,"reward":1.005},{"n":8,"raw":320.2,"best":321.6,"ok":true,"reward":1.013},{"n":9,"raw":316.0,"best":321.6,"ok":true,"reward":1.006},{"n":10,"raw":319.7,"best":321.6,"ok":true,"reward":1.012},{"n":11,"raw":317.5,"best":321.6,"ok":true,"reward":1.009}],"final_speedup":296.663588,"vs_oracle":0.950845},"gpt-5.6":{"oracle":312.0,"traj":[{"n":3,"raw":328.4,"best":328.4,"ok":true,"reward":1.026},{"n":4,"raw":325.5,"best":328.4,"ok":true,"reward":1.022},{"n":5,"raw":329.1,"best":329.1,"ok":true,"reward":1.027}],"final_speedup":322.002608,"vs_oracle":1.03206},"glm-5.2":{"oracle":312.0,"traj":[{"n":1,"raw":266.3,"best":266.3,"ok":true,"reward":0.9268},{"n":2,"raw":304.8,"best":304.8,"ok":true,"reward":0.9885},{"n":3,"raw":376.0,"best":376.0,"ok":true,"reward":1.102},{"n":4,"raw":399.1,"best":399.1,"ok":true,"reward":1.14},{"n":5,"raw":425.6,"best":425.6,"ok":true,"reward":1.182},{"n":6,"raw":429.9,"best":429.9,"ok":true,"reward":1.189},{"n":7,"raw":418.3,"best":429.9,"ok":true,"reward":1.17}],"final_speedup":428.981322,"vs_oracle":1.37494},"sonnet-5":{"oracle":312.0,"traj":[{"n":1,"raw":107.4,"best":107.4,"ok":true,"reward":0.6722},{"n":2,"raw":116.8,"best":116.8,"ok":true,"reward":0.6872},{"n":3,"raw":203.4,"best":203.4,"ok":true,"reward":0.826},{"n":4,"raw":215.6,"best":215.6,"ok":true,"reward":0.8456},{"n":5,"raw":217.4,"best":217.4,"ok":true,"reward":0.8485},{"n":6,"raw":218.9,"best":218.9,"ok":true,"reward":0.8508},{"n":7,"raw":220.4,"best":220.4,"ok":true,"reward":0.8533},{"n":8,"raw":223.5,"best":223.5,"ok":true,"reward":0.8581},{"n":9,"raw":217.1,"best":223.5,"ok":true,"reward":0.8479},{"n":10,"raw":219.0,"best":223.5,"ok":true,"reward":0.851},{"n":11,"raw":220.7,"best":223.5,"ok":true,"reward":0.8537},{"n":12,"raw":222.4,"best":223.5,"ok":true,"reward":0.8564}],"final_speedup":122.931915,"vs_oracle":0.394013},"qwen3.7":{"oracle":312.0,"traj":[{"n":1,"raw":250.2,"best":250.2,"ok":true,"reward":0.9009},{"n":2,"raw":285.3,"best":285.3,"ok":true,"reward":0.9573},{"n":3,"raw":311.2,"best":311.2,"ok":true,"reward":0.9987},{"n":4,"raw":305.1,"best":311.2,"ok":true,"reward":0.989},{"n":5,"raw":293.4,"best":311.2,"ok":true,"reward":0.9701},{"n":6,"raw":370.1,"best":370.1,"ok":true,"reward":1.093},{"n":7,"raw":386.8,"best":386.8,"ok":true,"reward":1.12},{"n":8,"raw":372.9,"best":386.8,"ok":true,"reward":1.098},{"n":9,"raw":406.0,"best":406.0,"ok":true,"reward":1.151},{"n":10,"raw":413.2,"best":413.2,"ok":true,"reward":1.162},{"n":11,"raw":425.8,"best":425.8,"ok":true,"reward":1.182},{"n":12,"raw":410.1,"best":425.8,"ok":true,"reward":1.157},{"n":13,"raw":386.6,"best":425.8,"ok":true,"reward":1.119},{"n":14,"raw":403.3,"best":425.8,"ok":true,"reward":1.146}],"final_speedup":419.959289,"vs_oracle":1.346023},"qwen3.8":{"oracle":312.0,"traj":[{"n":1,"raw":103.5,"best":103.5,"ok":true,"reward":0.6659},{"n":2,"raw":278.1,"best":278.1,"ok":true,"reward":0.9456},{"n":3,"raw":268.4,"best":278.1,"ok":true,"reward":0.9302},{"n":4,"raw":285.5,"best":285.5,"ok":true,"reward":0.9575},{"n":5,"raw":281.2,"best":285.5,"ok":true,"reward":0.9506},{"n":6,"raw":286.1,"best":286.1,"ok":true,"reward":0.9585},{"n":7,"raw":304.7,"best":304.7,"ok":true,"reward":0.9884},{"n":8,"raw":351.6,"best":351.6,"ok":true,"reward":1.063},{"n":9,"raw":354.5,"best":354.5,"ok":true,"reward":1.068},{"n":10,"raw":309.7,"best":354.5,"ok":true,"reward":0.9964},{"n":11,"raw":349.4,"best":354.5,"ok":true,"reward":1.06}],"final_speedup":347.104763,"vs_oracle":1.112515},"deepseek":{"oracle":312.0,"traj":[{"n":1,"raw":247.6,"best":247.6,"ok":true,"reward":0.8968},{"n":2,"raw":126.6,"best":247.6,"ok":true,"reward":0.7028},{"n":3,"raw":242.5,"best":247.6,"ok":true,"reward":0.8887},{"n":4,"raw":292.3,"best":292.3,"ok":true,"reward":0.9684},{"n":5,"raw":332.7,"best":332.7,"ok":true,"reward":1.033},{"n":6,"raw":336.5,"best":336.5,"ok":true,"reward":1.039},{"n":7,"raw":333.9,"best":336.5,"ok":true,"reward":1.035},{"n":8,"raw":379.3,"best":379.3,"ok":true,"reward":1.108}],"final_speedup":314.429246,"vs_oracle":1.007786}}},{"task_id":"wro-fla-nsa-sparse-sol","short_name":"NSA native sparse-attention kernel","infra_name":"native_sparse_attn_select","summary":"This is parallel_nsa, the selection path of Native Sparse Attention (NSA) in flash-linear-attention. The existing implementation is correct but slow; you must make it fast: under a causal mask, each query attends only to the keys within its own selected KV blocks. The input is q/k/v and the per-query selected block indices (grouped-query attention, G=HQ//H), and the output is the sparse-attention output. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 1036.42): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"C","medium_topic":"C2","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":1036.420558,"traj":[{"n":1,"raw":983.5,"best":983.5,"ok":true,"reward":0.9744},{"n":2,"raw":1214.0,"best":1214.0,"ok":true,"reward":1.086},{"n":3,"raw":1227.0,"best":1227.0,"ok":true,"reward":1.092},{"n":4,"raw":1269.0,"best":1269.0,"ok":true,"reward":1.112},{"n":5,"raw":1290.0,"best":1290.0,"ok":true,"reward":1.122},{"n":6,"raw":1279.0,"best":1290.0,"ok":true,"reward":1.117}],"final_speedup":1241.385315,"vs_oracle":1.197762},"kimi-k3":{"oracle":1036.420558,"traj":[{"n":1,"raw":821.1,"best":821.1,"ok":true,"reward":0.8961},{"n":2,"raw":1027.0,"best":1027.0,"ok":true,"reward":0.9953},{"n":3,"raw":1028.0,"best":1028.0,"ok":true,"reward":0.9959},{"n":4,"raw":1023.0,"best":1028.0,"ok":true,"reward":0.9935},{"n":5,"raw":1069.0,"best":1069.0,"ok":true,"reward":1.015}],"final_speedup":1079.572906,"vs_oracle":1.041636},"gpt-5.6":{"oracle":1036.420558,"traj":[{"n":4,"raw":1344.0,"best":1344.0,"ok":true,"reward":1.149},{"n":10,"raw":941.4,"best":1344.0,"ok":true,"reward":0.9542},{"n":11,"raw":1277.0,"best":1344.0,"ok":true,"reward":1.116}],"final_speedup":1270.268742,"vs_oracle":1.225631},"glm-5.2":{"oracle":1036.420558,"traj":[{"n":1,"raw":46.65,"best":46.65,"ok":true,"reward":0.5225},{"n":2,"raw":53.75,"best":53.75,"ok":true,"reward":0.5259},{"n":3,"raw":452.6,"best":452.6,"ok":true,"reward":0.7184},{"n":4,"raw":746.4,"best":746.4,"ok":true,"reward":0.8601},{"n":5,"raw":919.7,"best":919.7,"ok":true,"reward":0.9437},{"n":6,"raw":1048.0,"best":1048.0,"ok":true,"reward":1.006},{"n":7,"raw":912.2,"best":1048.0,"ok":true,"reward":0.9401},{"n":8,"raw":911.8,"best":1048.0,"ok":true,"reward":0.9399}],"final_speedup":932.024516,"vs_oracle":0.899273},"sonnet-5":{"oracle":1036.420558,"traj":[{"n":1,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":2,"raw":1110.0,"best":1110.0,"ok":true,"reward":1.035},{"n":3,"raw":1169.0,"best":1169.0,"ok":true,"reward":1.064},{"n":4,"raw":1352.0,"best":1352.0,"ok":true,"reward":1.152}],"final_speedup":1330.537312,"vs_oracle":1.283781},"qwen3.7":{"oracle":1036.420558,"traj":[{"n":1,"raw":25.39,"best":25.39,"ok":true,"reward":0.5122},{"n":2,"raw":24.86,"best":25.39,"ok":true,"reward":0.512},{"n":3,"raw":5.002,"best":25.39,"ok":true,"reward":0.5024},{"n":4,"raw":34.05,"best":34.05,"ok":true,"reward":0.5164},{"n":5,"raw":34.99,"best":34.99,"ok":true,"reward":0.5169},{"n":6,"raw":35.31,"best":35.31,"ok":true,"reward":0.517},{"n":7,"raw":37.13,"best":37.13,"ok":true,"reward":0.5179},{"n":8,"raw":37.98,"best":37.98,"ok":true,"reward":0.5183},{"n":9,"raw":37.08,"best":37.98,"ok":true,"reward":0.5179},{"n":10,"raw":51.01,"best":51.01,"ok":true,"reward":0.5246},{"n":11,"raw":34.12,"best":51.01,"ok":true,"reward":0.5165},{"n":12,"raw":52.07,"best":52.07,"ok":true,"reward":0.5251},{"n":13,"raw":49.48,"best":52.07,"ok":true,"reward":0.5239},{"n":14,"raw":50.57,"best":52.07,"ok":true,"reward":0.5244},{"n":15,"raw":50.98,"best":52.07,"ok":true,"reward":0.5246}],"final_speedup":51.791226,"vs_oracle":0.049971},"qwen3.8":{"oracle":1036.420558,"traj":[{"n":1,"raw":669.1,"best":669.1,"ok":true,"reward":0.8228},{"n":2,"raw":859.4,"best":859.4,"ok":true,"reward":0.9146},{"n":3,"raw":872.2,"best":872.2,"ok":true,"reward":0.9208}],"final_speedup":876.025728,"vs_oracle":0.845242},"deepseek":{"oracle":1036.420558,"traj":[{"n":1,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":2,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":3,"raw":0.0,"best":0.0,"ok":true,"reward":0.375},{"n":4,"raw":0.0,"best":0.0,"ok":true,"reward":0.375},{"n":5,"raw":0.0,"best":0.0,"ok":true,"reward":0.375},{"n":6,"raw":0.0,"best":0.0,"ok":true,"reward":0.375},{"n":7,"raw":0.0,"best":0.0,"ok":true,"reward":0.375},{"n":8,"raw":0.9886,"best":0.9886,"ok":true,"reward":0.5005},{"n":9,"raw":1.336,"best":1.336,"ok":true,"reward":0.5006},{"n":10,"raw":1.368,"best":1.368,"ok":true,"reward":0.5007},{"n":11,"raw":1.13,"best":1.368,"ok":true,"reward":0.5005},{"n":12,"raw":1.414,"best":1.414,"ok":true,"reward":0.5007},{"n":13,"raw":1.89,"best":1.89,"ok":true,"reward":0.5009},{"n":14,"raw":2.178,"best":2.178,"ok":true,"reward":0.5011},{"n":15,"raw":2.648,"best":2.648,"ok":true,"reward":0.5013},{"n":16,"raw":2.729,"best":2.729,"ok":true,"reward":0.5013}],"final_speedup":2.664001,"vs_oracle":0.00257}}},{"task_id":"wro-flashattn-cute-blocksparse-bwd","short_name":"FlashAttention CuTe block-sparse backward","infra_name":"blocksparse_attn_backward","summary":"This is the differentiable block-sparse multi-head attention in the flash-attention CuTe version: given the allowed (query-block, key-block) pairs, a query may only attend to keys within the allowed blocks. The existing implementation is correct but slow; you must make both the forward and the backward(dout) pass fast, and the signature may not be changed. The input is q/k/v and the block-sparse mask description, and the output is the attention result plus the backward dq/dk/dv. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 49.0085): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"F","medium_topic":"F3","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":49.008492,"traj":[{"n":1,"raw":39.85,"best":39.85,"ok":true,"reward":0.4734}],"final_speedup":49.717806,"vs_oracle":1.014473},"kimi-k3":{"oracle":49.008492,"traj":[{"n":1,"raw":45.42,"best":45.42,"ok":true,"reward":0.4902}],"final_speedup":45.191242,"vs_oracle":0.92211},"gpt-5.6":{"oracle":49.008492,"traj":[{"n":1,"raw":33.14,"best":33.14,"ok":true,"reward":0.4497}],"final_speedup":32.351892,"vs_oracle":0.660128},"glm-5.2":{"oracle":49.008492,"traj":[{"n":1,"raw":4.092,"best":4.092,"ok":true,"reward":0.181},{"n":2,"raw":5.132,"best":5.132,"ok":true,"reward":0.2101},{"n":3,"raw":5.727,"best":5.727,"ok":true,"reward":0.2242},{"n":4,"raw":5.977,"best":5.977,"ok":true,"reward":0.2297},{"n":5,"raw":6.014,"best":6.014,"ok":true,"reward":0.2305},{"n":6,"raw":7.355,"best":7.355,"ok":true,"reward":0.2563},{"n":7,"raw":8.338,"best":8.338,"ok":true,"reward":0.2725},{"n":8,"raw":8.365,"best":8.365,"ok":true,"reward":0.2729},{"n":9,"raw":11.07,"best":11.07,"ok":true,"reward":0.3089},{"n":10,"raw":13.75,"best":13.75,"ok":true,"reward":0.3367},{"n":11,"raw":13.99,"best":13.99,"ok":true,"reward":0.339},{"n":12,"raw":13.72,"best":13.99,"ok":true,"reward":0.3365}],"final_speedup":13.441629,"vs_oracle":0.274271},"sonnet-5":{"oracle":49.008492,"traj":[{"n":1,"raw":41.07,"best":41.07,"ok":true,"reward":0.4773}],"final_speedup":48.520177,"vs_oracle":0.990036},"qwen3.7":{"oracle":49.008492,"traj":[{"n":1,"raw":49.11,"best":49.11,"ok":true,"reward":0.5003}],"final_speedup":48.979942,"vs_oracle":0.999417},"qwen3.8":{"oracle":49.008492,"traj":[{"n":1,"raw":49.3,"best":49.3,"ok":true,"reward":0.5008},{"n":2,"raw":48.81,"best":49.3,"ok":true,"reward":0.4995}],"final_speedup":34.981289,"vs_oracle":0.71378},"deepseek":{"oracle":49.008492,"traj":[{"n":1,"raw":49.02,"best":49.02,"ok":true,"reward":0.5}],"final_speedup":48.985937,"vs_oracle":0.99954}}},{"task_id":"wro-causal-delivery-vclock-coupled","short_name":"Vector-clock causal broadcast delivery, indexed","infra_name":"causal_delivery_vclock","summary":"This is the delivery layer of a causal-broadcast messaging system: the receiver must hand messages to the application in happens-before order, buffering out-of-order arrivals first. The two coupled files under causal/ are correct but slow when there are many out-of-order arrivals; you must speed up the buffering and delivery decision. The input is an out-of-order message stream carrying vector clocks, and the output is a delivery order that satisfies causal order. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 187.501): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"I","medium_topic":"I2","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":187.500717,"traj":[{"n":1,"raw":701.6,"best":null,"ok":false,"reward":0.0},{"n":2,"raw":610.8,"best":null,"ok":false,"reward":0.0},{"n":3,"raw":531.2,"best":null,"ok":false,"reward":0.0},{"n":4,"raw":518.2,"best":null,"ok":false,"reward":0.0},{"n":5,"raw":1.099,"best":1.099,"ok":true,"reward":0.5029},{"n":6,"raw":487.8,"best":1.099,"ok":false,"reward":0.0},{"n":7,"raw":253.3,"best":253.3,"ok":true,"reward":1.175},{"n":8,"raw":297.1,"best":297.1,"ok":true,"reward":1.292},{"n":9,"raw":307.7,"best":307.7,"ok":true,"reward":1.321},{"n":10,"raw":335.9,"best":335.9,"ok":true,"reward":1.396},{"n":11,"raw":468.7,"best":335.9,"ok":false,"reward":0.0},{"n":12,"raw":444.5,"best":335.9,"ok":false,"reward":0.0},{"n":13,"raw":371.8,"best":371.8,"ok":true,"reward":1.492},{"n":14,"raw":389.5,"best":371.8,"ok":false,"reward":0.0}],"final_speedup":380.064487,"vs_oracle":2.027003},"kimi-k3":{"oracle":187.500717,"traj":[{"n":1,"raw":404.2,"best":null,"ok":false,"reward":0.0},{"n":2,"raw":404.1,"best":null,"ok":false,"reward":0.0},{"n":3,"raw":412.1,"best":null,"ok":false,"reward":0.0},{"n":4,"raw":365.0,"best":365.0,"ok":true,"reward":1.473},{"n":5,"raw":382.3,"best":365.0,"ok":false,"reward":0.0},{"n":6,"raw":357.5,"best":365.0,"ok":true,"reward":1.453},{"n":7,"raw":377.3,"best":365.0,"ok":false,"reward":0.0},{"n":8,"raw":362.8,"best":365.0,"ok":true,"reward":1.468},{"n":9,"raw":365.5,"best":365.5,"ok":true,"reward":1.475},{"n":10,"raw":390.2,"best":365.5,"ok":false,"reward":0.0},{"n":11,"raw":363.4,"best":365.5,"ok":true,"reward":1.469},{"n":12,"raw":367.4,"best":367.4,"ok":true,"reward":1.48},{"n":13,"raw":394.1,"best":367.4,"ok":false,"reward":0.0}],"final_speedup":364.029188,"vs_oracle":1.941482},"gpt-5.6":{"oracle":187.500717,"traj":[{"n":7,"raw":368.3,"best":368.3,"ok":true,"reward":1.482}],"final_speedup":385.519112,"vs_oracle":2.056094},"glm-5.2":{"oracle":187.500717,"traj":[{"n":1,"raw":487.4,"best":null,"ok":false,"reward":0.0},{"n":2,"raw":587.9,"best":null,"ok":false,"reward":0.0},{"n":3,"raw":524.3,"best":null,"ok":false,"reward":0.0},{"n":4,"raw":666.2,"best":null,"ok":false,"reward":0.0},{"n":5,"raw":574.0,"best":null,"ok":false,"reward":0.0},{"n":6,"raw":614.8,"best":null,"ok":false,"reward":0.0},{"n":7,"raw":0.9574,"best":0.9574,"ok":true,"reward":0.5026},{"n":8,"raw":579.5,"best":0.9574,"ok":false,"reward":0.0},{"n":9,"raw":16.47,"best":16.47,"ok":true,"reward":0.5439},{"n":10,"raw":207.1,"best":207.1,"ok":true,"reward":1.052},{"n":11,"raw":647.2,"best":207.1,"ok":false,"reward":0.0},{"n":12,"raw":561.1,"best":207.1,"ok":false,"reward":0.0},{"n":13,"raw":393.5,"best":207.1,"ok":false,"reward":0.0},{"n":14,"raw":404.0,"best":207.1,"ok":false,"reward":0.0},{"n":15,"raw":230.2,"best":230.2,"ok":true,"reward":1.114},{"n":16,"raw":217.0,"best":230.2,"ok":true,"reward":1.079}],"final_speedup":236.876587,"vs_oracle":1.263337},"sonnet-5":{"oracle":187.500717,"traj":[{"n":1,"raw":343.7,"best":343.7,"ok":true,"reward":1.417},{"n":2,"raw":447.1,"best":343.7,"ok":false,"reward":0.0},{"n":3,"raw":414.5,"best":343.7,"ok":false,"reward":0.0},{"n":4,"raw":412.1,"best":343.7,"ok":false,"reward":0.0},{"n":5,"raw":229.1,"best":343.7,"ok":true,"reward":1.111},{"n":6,"raw":258.5,"best":343.7,"ok":true,"reward":1.189},{"n":7,"raw":259.8,"best":343.7,"ok":true,"reward":1.193},{"n":8,"raw":269.1,"best":343.7,"ok":true,"reward":1.218},{"n":9,"raw":423.1,"best":343.7,"ok":false,"reward":0.0},{"n":10,"raw":322.7,"best":343.7,"ok":true,"reward":1.36},{"n":11,"raw":394.4,"best":343.7,"ok":false,"reward":0.0},{"n":12,"raw":320.6,"best":343.7,"ok":true,"reward":1.355},{"n":13,"raw":318.8,"best":343.7,"ok":true,"reward":1.35},{"n":14,"raw":340.5,"best":343.7,"ok":true,"reward":1.408}],"final_speedup":350.068679,"vs_oracle":1.867026},"qwen3.7":{"oracle":187.500717,"traj":[{"n":1,"raw":0.0,"best":null,"ok":false,"reward":0.3269},{"n":2,"raw":164.3,"best":164.3,"ok":true,"reward":0.938},{"n":3,"raw":206.5,"best":206.5,"ok":true,"reward":1.051},{"n":4,"raw":0.0,"best":206.5,"ok":false,"reward":0.0},{"n":5,"raw":336.2,"best":336.2,"ok":true,"reward":1.397},{"n":6,"raw":386.3,"best":336.2,"ok":false,"reward":0.0},{"n":7,"raw":381.4,"best":336.2,"ok":false,"reward":0.0},{"n":8,"raw":374.8,"best":374.8,"ok":true,"reward":1.5},{"n":9,"raw":396.6,"best":374.8,"ok":false,"reward":0.0},{"n":10,"raw":399.5,"best":374.8,"ok":false,"reward":0.0},{"n":11,"raw":370.2,"best":374.8,"ok":true,"reward":1.487},{"n":12,"raw":377.7,"best":374.8,"ok":false,"reward":0.0},{"n":13,"raw":385.0,"best":374.8,"ok":false,"reward":0.0},{"n":14,"raw":387.9,"best":374.8,"ok":false,"reward":0.0},{"n":15,"raw":377.4,"best":374.8,"ok":false,"reward":0.0}],"final_speedup":369.958972,"vs_oracle":1.973107},"qwen3.8":{"oracle":187.500717,"traj":[{"n":1,"raw":477.0,"best":null,"ok":false,"reward":0.0},{"n":2,"raw":477.3,"best":null,"ok":false,"reward":0.0},{"n":3,"raw":1.044,"best":1.044,"ok":true,"reward":0.5028},{"n":4,"raw":0.1185,"best":1.044,"ok":true,"reward":0.5003},{"n":5,"raw":423.9,"best":1.044,"ok":false,"reward":0.0},{"n":6,"raw":371.5,"best":371.5,"ok":true,"reward":1.491},{"n":7,"raw":387.3,"best":371.5,"ok":false,"reward":0.0},{"n":8,"raw":378.9,"best":371.5,"ok":false,"reward":0.0},{"n":9,"raw":373.9,"best":373.9,"ok":true,"reward":1.497},{"n":10,"raw":367.2,"best":373.9,"ok":true,"reward":1.479}],"final_speedup":359.674692,"vs_oracle":1.918258},"deepseek":{"oracle":187.500717,"traj":[{"n":1,"raw":280.4,"best":280.4,"ok":true,"reward":1.248},{"n":2,"raw":195.3,"best":280.4,"ok":true,"reward":1.021},{"n":3,"raw":0.0,"best":280.4,"ok":false,"reward":0.0},{"n":4,"raw":368.6,"best":368.6,"ok":true,"reward":1.483},{"n":5,"raw":430.0,"best":368.6,"ok":false,"reward":0.0},{"n":6,"raw":441.8,"best":368.6,"ok":false,"reward":0.0},{"n":7,"raw":424.8,"best":368.6,"ok":false,"reward":0.0},{"n":8,"raw":420.6,"best":368.6,"ok":false,"reward":0.0},{"n":9,"raw":233.9,"best":368.6,"ok":true,"reward":1.124},{"n":10,"raw":0.0,"best":368.6,"ok":false,"reward":0.0},{"n":11,"raw":382.4,"best":368.6,"ok":false,"reward":0.0},{"n":12,"raw":363.0,"best":368.6,"ok":true,"reward":1.468},{"n":13,"raw":332.5,"best":368.6,"ok":true,"reward":1.387},{"n":14,"raw":376.5,"best":368.6,"ok":false,"reward":0.0},{"n":15,"raw":351.1,"best":368.6,"ok":true,"reward":1.436}],"final_speedup":356.363747,"vs_oracle":1.900599}}},{"task_id":"wro-secure-agg-shamir-committee-coupled","short_name":"Shamir secret-sharing secure aggregation","infra_name":"shamir_secure_aggregation","summary":"This is the server side of a privacy-preserving (secure aggregation) system: clients mask their model updates and split them across a committee using Shamir secret sharing, and the server reconstructs the aggregated value from these shares. The two coupled files under secureagg/ are correct but slow when reconstructing many coordinates over the same committee (every model coordinate must be reconstructed once per round); you must make it fast. The input is the per-coordinate lists of shares (over multiple rounds); the output is the reconstructed aggregated value. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 57.6582): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"I","medium_topic":"I4","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":57.658178,"traj":[{"n":1,"raw":169.9,"best":null,"ok":false,"reward":0.0},{"n":2,"raw":160.6,"best":null,"ok":false,"reward":0.0},{"n":3,"raw":130.9,"best":null,"ok":false,"reward":0.0},{"n":4,"raw":1.005,"best":1.005,"ok":true,"reward":0.5087},{"n":5,"raw":102.7,"best":102.7,"ok":true,"reward":1.39},{"n":6,"raw":200.2,"best":102.7,"ok":false,"reward":0.0},{"n":7,"raw":173.2,"best":102.7,"ok":false,"reward":0.0},{"n":8,"raw":172.3,"best":102.7,"ok":false,"reward":0.0},{"n":9,"raw":93.74,"best":102.7,"ok":true,"reward":1.313},{"n":10,"raw":190.0,"best":102.7,"ok":false,"reward":0.0},{"n":11,"raw":186.8,"best":102.7,"ok":false,"reward":0.0},{"n":12,"raw":113.4,"best":113.4,"ok":true,"reward":1.483},{"n":13,"raw":122.4,"best":113.4,"ok":false,"reward":0.0}],"final_speedup":112.904192,"vs_oracle":1.958164},"kimi-k3":{"oracle":57.658178,"traj":[{"n":1,"raw":108.4,"best":108.4,"ok":true,"reward":1.44},{"n":2,"raw":154.1,"best":108.4,"ok":false,"reward":0.0},{"n":3,"raw":116.1,"best":108.4,"ok":false,"reward":0.0},{"n":4,"raw":109.4,"best":109.4,"ok":true,"reward":1.449},{"n":5,"raw":106.5,"best":109.4,"ok":true,"reward":1.423},{"n":6,"raw":100.8,"best":109.4,"ok":true,"reward":1.374},{"n":7,"raw":104.2,"best":109.4,"ok":true,"reward":1.404},{"n":8,"raw":105.1,"best":109.4,"ok":true,"reward":1.411},{"n":9,"raw":118.8,"best":109.4,"ok":false,"reward":0.0},{"n":10,"raw":114.5,"best":114.5,"ok":true,"reward":1.493},{"n":11,"raw":145.4,"best":114.5,"ok":false,"reward":0.0},{"n":12,"raw":116.7,"best":114.5,"ok":false,"reward":0.0},{"n":13,"raw":112.5,"best":114.5,"ok":true,"reward":1.475},{"n":14,"raw":112.7,"best":114.5,"ok":true,"reward":1.477},{"n":15,"raw":113.7,"best":114.5,"ok":true,"reward":1.486},{"n":16,"raw":118.3,"best":114.5,"ok":false,"reward":0.0}],"final_speedup":114.706708,"vs_oracle":1.989427},"gpt-5.6":{"oracle":57.658178,"traj":[{"n":14,"raw":113.1,"best":113.1,"ok":true,"reward":1.481},{"n":15,"raw":115.3,"best":115.3,"ok":true,"reward":1.5}],"final_speedup":117.090043,"vs_oracle":2.030762},"glm-5.2":{"oracle":57.658178,"traj":[{"n":1,"raw":180.1,"best":null,"ok":false,"reward":0.0},{"n":2,"raw":170.9,"best":null,"ok":false,"reward":0.0},{"n":3,"raw":176.1,"best":null,"ok":false,"reward":0.0},{"n":4,"raw":175.3,"best":null,"ok":false,"reward":0.0},{"n":5,"raw":129.6,"best":null,"ok":false,"reward":0.0},{"n":6,"raw":129.0,"best":null,"ok":false,"reward":0.0},{"n":7,"raw":3.083,"best":3.083,"ok":true,"reward":0.5267},{"n":8,"raw":64.44,"best":64.44,"ok":true,"reward":1.059},{"n":9,"raw":98.76,"best":98.76,"ok":true,"reward":1.356},{"n":10,"raw":109.7,"best":109.7,"ok":true,"reward":1.452},{"n":11,"raw":123.7,"best":109.7,"ok":false,"reward":0.0},{"n":12,"raw":119.6,"best":109.7,"ok":false,"reward":0.0},{"n":13,"raw":119.4,"best":109.7,"ok":false,"reward":0.0},{"n":14,"raw":116.8,"best":109.7,"ok":false,"reward":0.0}],"final_speedup":109.974058,"vs_oracle":1.907345},"sonnet-5":{"oracle":57.658178,"traj":[{"n":1,"raw":59.15,"best":59.15,"ok":true,"reward":1.013},{"n":2,"raw":75.6,"best":75.6,"ok":true,"reward":1.156},{"n":3,"raw":85.67,"best":85.67,"ok":true,"reward":1.243},{"n":4,"raw":101.6,"best":101.6,"ok":true,"reward":1.381},{"n":5,"raw":134.0,"best":101.6,"ok":false,"reward":0.0},{"n":6,"raw":134.4,"best":101.6,"ok":false,"reward":0.0},{"n":7,"raw":101.2,"best":101.6,"ok":true,"reward":1.378},{"n":8,"raw":103.1,"best":103.1,"ok":true,"reward":1.394},{"n":9,"raw":104.3,"best":104.3,"ok":true,"reward":1.404}],"final_speedup":102.043156,"vs_oracle":1.769795},"qwen3.7":{"oracle":57.658178,"traj":[{"n":1,"raw":87.33,"best":87.33,"ok":true,"reward":1.257},{"n":2,"raw":175.4,"best":87.33,"ok":false,"reward":0.0},{"n":3,"raw":135.9,"best":87.33,"ok":false,"reward":0.0},{"n":4,"raw":139.6,"best":87.33,"ok":false,"reward":0.0},{"n":5,"raw":87.46,"best":87.46,"ok":true,"reward":1.258},{"n":6,"raw":88.07,"best":88.07,"ok":true,"reward":1.264},{"n":7,"raw":137.7,"best":88.07,"ok":false,"reward":0.0},{"n":8,"raw":116.2,"best":88.07,"ok":false,"reward":0.0},{"n":9,"raw":102.7,"best":102.7,"ok":true,"reward":1.39},{"n":10,"raw":180.1,"best":102.7,"ok":false,"reward":0.0},{"n":11,"raw":100.4,"best":102.7,"ok":true,"reward":1.371},{"n":12,"raw":101.4,"best":102.7,"ok":true,"reward":1.379},{"n":13,"raw":90.76,"best":102.7,"ok":true,"reward":1.287},{"n":14,"raw":95.86,"best":102.7,"ok":true,"reward":1.331},{"n":15,"raw":107.4,"best":107.4,"ok":true,"reward":1.431},{"n":16,"raw":106.3,"best":107.4,"ok":true,"reward":1.422}],"final_speedup":104.916295,"vs_oracle":1.819626},"qwen3.8":{"oracle":57.658178,"traj":[{"n":1,"raw":79.09,"best":79.09,"ok":true,"reward":1.186},{"n":2,"raw":109.1,"best":109.1,"ok":true,"reward":1.446},{"n":3,"raw":143.3,"best":109.1,"ok":false,"reward":0.0},{"n":4,"raw":148.2,"best":109.1,"ok":false,"reward":0.0},{"n":5,"raw":118.8,"best":109.1,"ok":false,"reward":0.0},{"n":6,"raw":105.7,"best":109.1,"ok":true,"reward":1.417},{"n":7,"raw":126.7,"best":109.1,"ok":false,"reward":0.0},{"n":8,"raw":109.3,"best":109.3,"ok":true,"reward":1.448}],"final_speedup":108.516827,"vs_oracle":1.882072},"deepseek":{"oracle":57.658178,"traj":[{"n":1,"raw":96.87,"best":96.87,"ok":true,"reward":1.34},{"n":2,"raw":179.9,"best":96.87,"ok":false,"reward":0.0},{"n":3,"raw":181.3,"best":96.87,"ok":false,"reward":0.0},{"n":4,"raw":173.2,"best":96.87,"ok":false,"reward":0.0},{"n":5,"raw":156.6,"best":96.87,"ok":false,"reward":0.0},{"n":6,"raw":0.9998,"best":96.87,"ok":true,"reward":0.5087},{"n":7,"raw":175.0,"best":96.87,"ok":false,"reward":0.0},{"n":8,"raw":153.9,"best":96.87,"ok":false,"reward":0.0},{"n":9,"raw":89.86,"best":96.87,"ok":true,"reward":1.279},{"n":10,"raw":96.27,"best":96.87,"ok":true,"reward":1.335},{"n":11,"raw":98.68,"best":98.68,"ok":true,"reward":1.356},{"n":12,"raw":148.7,"best":98.68,"ok":false,"reward":0.0},{"n":13,"raw":124.9,"best":98.68,"ok":false,"reward":0.0},{"n":14,"raw":98.03,"best":98.68,"ok":true,"reward":1.35},{"n":15,"raw":104.0,"best":104.0,"ok":true,"reward":1.402},{"n":16,"raw":93.22,"best":104.0,"ok":true,"reward":1.308}],"final_speedup":103.656608,"vs_oracle":1.797778}}},{"task_id":"wro-memory-accounting-sim-coupled","short_name":"GPU-memory accounting simulation","infra_name":"memory_plan_peak_accounting","summary":"This is a memory planner and OOM predictor for training/inference execution plans: each tensor is allocated at some step and freed at another, occupying nbytes while it is live. The two coupled files under memsim/ are correct but slow on large plans and under repeated peak queries; you must make the step-by-step usage, peak statistics, and eviction-schedule search fast. The inputs are each tensor's [alloc, free) lifetime, its byte size, and a budget; the outputs are the step-by-step memory footprint, the peak, and an eviction plan that fits within the budget. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 1344.4): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"H","medium_topic":"H2","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":1344.398352,"traj":[{"n":1,"raw":108.9,"best":108.9,"ok":true,"reward":0.3256},{"n":2,"raw":189.3,"best":189.3,"ok":true,"reward":0.3639},{"n":3,"raw":182.6,"best":189.3,"ok":true,"reward":0.3614},{"n":4,"raw":193.1,"best":193.1,"ok":true,"reward":0.3653},{"n":5,"raw":224.9,"best":224.9,"ok":true,"reward":0.3759},{"n":6,"raw":225.9,"best":225.9,"ok":true,"reward":0.3762}],"final_speedup":222.441169,"vs_oracle":0.165458},"kimi-k3":{"oracle":1344.398352,"traj":[{"n":1,"raw":71.39,"best":71.39,"ok":true,"reward":0.2962},{"n":2,"raw":72.39,"best":72.39,"ok":true,"reward":0.2972},{"n":3,"raw":72.73,"best":72.73,"ok":true,"reward":0.2975},{"n":4,"raw":72.5,"best":72.73,"ok":true,"reward":0.2973},{"n":5,"raw":81.57,"best":81.57,"ok":true,"reward":0.3055},{"n":6,"raw":81.12,"best":81.57,"ok":true,"reward":0.3051},{"n":7,"raw":82.83,"best":82.83,"ok":true,"reward":0.3066},{"n":8,"raw":82.26,"best":82.83,"ok":true,"reward":0.3061},{"n":9,"raw":81.76,"best":82.83,"ok":true,"reward":0.3057},{"n":10,"raw":82.97,"best":82.97,"ok":true,"reward":0.3067},{"n":11,"raw":82.1,"best":82.97,"ok":true,"reward":0.3059},{"n":12,"raw":82.26,"best":82.97,"ok":true,"reward":0.3061},{"n":13,"raw":83.9,"best":83.9,"ok":true,"reward":0.3075},{"n":14,"raw":83.96,"best":83.96,"ok":true,"reward":0.3075},{"n":15,"raw":83.87,"best":83.96,"ok":true,"reward":0.3074},{"n":16,"raw":83.9,"best":83.96,"ok":true,"reward":0.3075}],"final_speedup":85.496903,"vs_oracle":0.063595},"gpt-5.6":{"oracle":1344.398352,"traj":[{"n":12,"raw":117.6,"best":117.6,"ok":true,"reward":0.3309},{"n":14,"raw":116.6,"best":117.6,"ok":true,"reward":0.3303}],"final_speedup":112.272763,"vs_oracle":0.083512},"glm-5.2":{"oracle":1344.398352,"traj":[{"n":1,"raw":1157.0,"best":1157.0,"ok":true,"reward":0.4896},{"n":2,"raw":1138.0,"best":1157.0,"ok":true,"reward":0.4884},{"n":3,"raw":1156.0,"best":1157.0,"ok":true,"reward":0.4895},{"n":4,"raw":1353.0,"best":1353.0,"ok":true,"reward":0.5004},{"n":5,"raw":1377.0,"best":1377.0,"ok":true,"reward":0.5016},{"n":6,"raw":1392.0,"best":1392.0,"ok":true,"reward":0.5024},{"n":7,"raw":1402.0,"best":1402.0,"ok":true,"reward":0.5029},{"n":8,"raw":1418.0,"best":1418.0,"ok":true,"reward":0.5037},{"n":9,"raw":1454.0,"best":1454.0,"ok":true,"reward":0.5055},{"n":10,"raw":1455.0,"best":1455.0,"ok":true,"reward":0.5055},{"n":11,"raw":1670.0,"best":1670.0,"ok":true,"reward":0.5151},{"n":12,"raw":1737.0,"best":1737.0,"ok":true,"reward":0.5178},{"n":13,"raw":1734.0,"best":1737.0,"ok":true,"reward":0.5177},{"n":14,"raw":1755.0,"best":1755.0,"ok":true,"reward":0.5185}],"final_speedup":1671.136977,"vs_oracle":1.243037},"sonnet-5":{"oracle":1344.398352,"traj":[{"n":1,"raw":104.8,"best":104.8,"ok":true,"reward":0.3229},{"n":2,"raw":106.7,"best":106.7,"ok":true,"reward":0.3241},{"n":3,"raw":109.0,"best":109.0,"ok":true,"reward":0.3256},{"n":4,"raw":110.0,"best":110.0,"ok":true,"reward":0.3263},{"n":5,"raw":105.0,"best":110.0,"ok":true,"reward":0.3231},{"n":6,"raw":107.2,"best":110.0,"ok":true,"reward":0.3245},{"n":7,"raw":108.8,"best":110.0,"ok":true,"reward":0.3255},{"n":8,"raw":106.5,"best":110.0,"ok":true,"reward":0.324},{"n":9,"raw":106.7,"best":110.0,"ok":true,"reward":0.3241},{"n":10,"raw":106.8,"best":110.0,"ok":true,"reward":0.3242},{"n":11,"raw":105.7,"best":110.0,"ok":true,"reward":0.3235},{"n":12,"raw":108.2,"best":110.0,"ok":true,"reward":0.3251},{"n":13,"raw":108.2,"best":110.0,"ok":true,"reward":0.3251},{"n":14,"raw":106.1,"best":110.0,"ok":true,"reward":0.3238}],"final_speedup":107.160689,"vs_oracle":0.079709},"qwen3.7":{"oracle":1344.398352,"traj":[{"n":1,"raw":1125.0,"best":1125.0,"ok":true,"reward":0.4876},{"n":2,"raw":1151.0,"best":1151.0,"ok":true,"reward":0.4892},{"n":3,"raw":1148.0,"best":1151.0,"ok":true,"reward":0.489},{"n":4,"raw":1223.0,"best":1223.0,"ok":true,"reward":0.4934},{"n":5,"raw":1238.0,"best":1238.0,"ok":true,"reward":0.4943},{"n":6,"raw":1252.0,"best":1252.0,"ok":true,"reward":0.4951},{"n":7,"raw":1317.0,"best":1317.0,"ok":true,"reward":0.4986},{"n":8,"raw":1336.0,"best":1336.0,"ok":true,"reward":0.4996},{"n":9,"raw":1311.0,"best":1336.0,"ok":true,"reward":0.4983},{"n":10,"raw":1311.0,"best":1336.0,"ok":true,"reward":0.4983},{"n":11,"raw":1338.0,"best":1338.0,"ok":true,"reward":0.4997},{"n":12,"raw":1100.0,"best":1338.0,"ok":true,"reward":0.4861},{"n":13,"raw":1332.0,"best":1338.0,"ok":true,"reward":0.4994},{"n":14,"raw":1328.0,"best":1338.0,"ok":true,"reward":0.4991},{"n":15,"raw":1347.0,"best":1347.0,"ok":true,"reward":0.5002}],"final_speedup":1341.78676,"vs_oracle":0.998057},"qwen3.8":{"oracle":1344.398352,"traj":[{"n":1,"raw":1274.0,"best":1274.0,"ok":true,"reward":0.4963},{"n":2,"raw":1275.0,"best":1275.0,"ok":true,"reward":0.4963},{"n":3,"raw":1398.0,"best":1398.0,"ok":true,"reward":0.5027},{"n":4,"raw":1361.0,"best":1398.0,"ok":true,"reward":0.5009},{"n":5,"raw":1400.0,"best":1400.0,"ok":true,"reward":0.5028},{"n":6,"raw":1389.0,"best":1400.0,"ok":true,"reward":0.5022},{"n":7,"raw":1438.0,"best":1438.0,"ok":true,"reward":0.5047},{"n":8,"raw":1470.0,"best":1470.0,"ok":true,"reward":0.5062},{"n":9,"raw":1422.0,"best":1470.0,"ok":true,"reward":0.5039},{"n":10,"raw":1428.0,"best":1470.0,"ok":true,"reward":0.5042},{"n":11,"raw":1429.0,"best":1470.0,"ok":true,"reward":0.5042}],"final_speedup":1459.009393,"vs_oracle":1.085251},"deepseek":{"oracle":1344.398352,"traj":[{"n":1,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":2,"raw":1138.0,"best":1138.0,"ok":true,"reward":0.4884},{"n":3,"raw":1129.0,"best":1138.0,"ok":true,"reward":0.4879},{"n":4,"raw":1161.0,"best":1161.0,"ok":true,"reward":0.4898},{"n":5,"raw":1223.0,"best":1223.0,"ok":true,"reward":0.4934},{"n":6,"raw":1255.0,"best":1255.0,"ok":true,"reward":0.4952},{"n":7,"raw":1243.0,"best":1255.0,"ok":true,"reward":0.4946},{"n":8,"raw":1303.0,"best":1303.0,"ok":true,"reward":0.4978},{"n":9,"raw":1182.0,"best":1303.0,"ok":true,"reward":0.4911},{"n":10,"raw":1260.0,"best":1303.0,"ok":true,"reward":0.4955},{"n":11,"raw":1261.0,"best":1303.0,"ok":true,"reward":0.4956},{"n":12,"raw":1317.0,"best":1317.0,"ok":true,"reward":0.4986},{"n":13,"raw":1300.0,"best":1317.0,"ok":true,"reward":0.4977},{"n":14,"raw":1305.0,"best":1317.0,"ok":true,"reward":0.4979},{"n":15,"raw":1294.0,"best":1317.0,"ok":true,"reward":0.4973},{"n":16,"raw":1268.0,"best":1317.0,"ok":true,"reward":0.4959}],"final_speedup":1292.834042,"vs_oracle":0.961645}}},{"task_id":"wli-torchtitan-gptoss-expert-compute","short_name":"wli-torchtitan-gptoss-expert-compute","infra_name":"gptoss_moe_expert_compute","summary":"This is the expert computation of the gpt-oss MoE layer in torchtitan (PyTorch's native LLM training platform), with two coupled files on the per-token forward path. You must fix the routing counts and the grouped SwiGLU expert computation: build the one-hot routing map and reduce it into per-expert token counts, then run gpt-oss's interleaved gated swiglu. The input is the token hidden states and the routing result, and the output is the hidden states after expert computation; note that gate/linear are the even/odd interleaved channels of the last dimension. Metric: binary implementation-class score — 1.0 only if all hidden cases and hard gates pass; any single case failure or a triggered forbidden-edit/cheat condition is 0.0, with no partial credit.","bench":"lh","big_topic":"A","medium_topic":"A3","task_class":"implementation","is_performance":false,"metric":"speedup","lower_better":false,"models":{"qwen3.8":{"oracle":null,"traj":[{"n":1,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":2,"raw":0.0,"best":0.0,"ok":true,"reward":1.0}],"final_speedup":null,"vs_oracle":null}}},{"task_id":"wro-fla-wallattn-flash-sol","short_name":"WallAttention decay-window flash-attn kernel","infra_name":"windowed_decay_attention","summary":"This is parallel_wall_attn, the sliding-window decayed attention in flash-linear-attention, with per-channel multiplicative decay. The existing implementation is correct but slow; you must make it fast (the logit contains a difference of the log-domain prefix P=cumsum(g)/ln2). The input is q/k/v and the per-channel log decay g, and the output is the causal sliding-window attention result. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 308.864): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"D","medium_topic":"D1","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":308.863663,"traj":[{"n":1,"raw":328.5,"best":328.5,"ok":true,"reward":1.032},{"n":2,"raw":382.2,"best":382.2,"ok":true,"reward":1.119},{"n":3,"raw":380.0,"best":382.2,"ok":true,"reward":1.115},{"n":4,"raw":500.0,"best":500.0,"ok":true,"reward":1.309},{"n":5,"raw":496.1,"best":500.0,"ok":true,"reward":1.303},{"n":6,"raw":513.5,"best":513.5,"ok":true,"reward":1.331},{"n":7,"raw":522.2,"best":522.2,"ok":true,"reward":1.345},{"n":8,"raw":542.9,"best":542.9,"ok":true,"reward":1.379},{"n":9,"raw":601.0,"best":601.0,"ok":true,"reward":1.473},{"n":10,"raw":603.3,"best":603.3,"ok":true,"reward":1.477},{"n":11,"raw":605.8,"best":605.8,"ok":true,"reward":1.481}],"final_speedup":606.135387,"vs_oracle":1.962469},"kimi-k3":{"oracle":null,"traj":[{"n":1,"raw":215.3,"best":215.3,"ok":true,"reward":0.8486},{"n":2,"raw":198.2,"best":215.3,"ok":true,"reward":0.8209},{"n":3,"raw":202.3,"best":215.3,"ok":true,"reward":0.8275}],"final_speedup":233.361227,"vs_oracle":0.755548},"gpt-5.6":{"oracle":308.863663,"traj":[{"n":11,"raw":410.5,"best":410.5,"ok":true,"reward":1.165},{"n":14,"raw":407.8,"best":410.5,"ok":true,"reward":1.16}],"final_speedup":404.585567,"vs_oracle":1.309916},"glm-5.2":{"oracle":308.863663,"traj":[{"n":1,"raw":215.7,"best":215.7,"ok":true,"reward":0.8493},{"n":2,"raw":338.6,"best":338.6,"ok":true,"reward":1.048},{"n":3,"raw":312.8,"best":338.6,"ok":true,"reward":1.006},{"n":4,"raw":364.5,"best":364.5,"ok":true,"reward":1.09},{"n":5,"raw":366.3,"best":366.3,"ok":true,"reward":1.093}],"final_speedup":359.660585,"vs_oracle":1.164464},"sonnet-5":{"oracle":308.863663,"traj":[{"n":1,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":2,"raw":269.1,"best":269.1,"ok":true,"reward":0.9356},{"n":3,"raw":267.7,"best":269.1,"ok":true,"reward":0.9334},{"n":4,"raw":266.9,"best":269.1,"ok":true,"reward":0.932}],"final_speedup":267.794991,"vs_oracle":0.867033},"qwen3.7":{"oracle":308.863663,"traj":[{"n":1,"raw":54.91,"best":54.91,"ok":true,"reward":0.5889},{"n":2,"raw":39.63,"best":54.91,"ok":true,"reward":0.5642},{"n":3,"raw":40.89,"best":54.91,"ok":true,"reward":0.5662},{"n":4,"raw":58.8,"best":58.8,"ok":true,"reward":0.5952},{"n":5,"raw":0.0,"best":58.8,"ok":true,"reward":0.3333},{"n":6,"raw":56.9,"best":58.8,"ok":true,"reward":0.5921},{"n":7,"raw":57.55,"best":58.8,"ok":true,"reward":0.5932},{"n":8,"raw":59.06,"best":59.06,"ok":true,"reward":0.5956},{"n":9,"raw":71.36,"best":71.36,"ok":true,"reward":0.6155},{"n":10,"raw":74.26,"best":74.26,"ok":true,"reward":0.6202},{"n":11,"raw":61.07,"best":74.26,"ok":true,"reward":0.5989},{"n":12,"raw":70.79,"best":74.26,"ok":true,"reward":0.6146},{"n":13,"raw":72.76,"best":74.26,"ok":true,"reward":0.6178},{"n":14,"raw":73.59,"best":74.26,"ok":true,"reward":0.6191},{"n":15,"raw":0.7986,"best":74.26,"ok":true,"reward":0.5013}],"final_speedup":74.336061,"vs_oracle":0.240676},"qwen3.8":{"oracle":308.863663,"traj":[{"n":1,"raw":220.6,"best":220.6,"ok":true,"reward":0.8572},{"n":2,"raw":249.8,"best":249.8,"ok":true,"reward":0.9043},{"n":3,"raw":285.5,"best":285.5,"ok":true,"reward":0.9621},{"n":4,"raw":377.3,"best":377.3,"ok":true,"reward":1.111},{"n":5,"raw":430.4,"best":430.4,"ok":true,"reward":1.197},{"n":6,"raw":500.4,"best":500.4,"ok":true,"reward":1.31},{"n":7,"raw":526.9,"best":526.9,"ok":true,"reward":1.353},{"n":8,"raw":515.6,"best":526.9,"ok":true,"reward":1.335},{"n":9,"raw":545.0,"best":545.0,"ok":true,"reward":1.382},{"n":10,"raw":565.0,"best":565.0,"ok":true,"reward":1.415},{"n":11,"raw":577.2,"best":577.2,"ok":true,"reward":1.434}],"final_speedup":570.831433,"vs_oracle":1.848166},"deepseek":{"oracle":null,"traj":[{"n":1,"raw":1.002,"best":1.002,"ok":true,"reward":0.5016},{"n":2,"raw":1.0,"best":1.002,"ok":true,"reward":0.5016},{"n":3,"raw":0.9997,"best":1.002,"ok":true,"reward":0.5016},{"n":4,"raw":0.9957,"best":1.002,"ok":true,"reward":0.5016},{"n":5,"raw":1.0,"best":1.002,"ok":true,"reward":0.5016},{"n":6,"raw":37.07,"best":37.07,"ok":true,"reward":0.56},{"n":7,"raw":46.9,"best":46.9,"ok":true,"reward":0.5759},{"n":8,"raw":0.0,"best":46.9,"ok":true,"reward":0.0},{"n":9,"raw":46.89,"best":46.9,"ok":true,"reward":0.5759},{"n":10,"raw":45.88,"best":46.9,"ok":true,"reward":0.5743},{"n":11,"raw":234.8,"best":234.8,"ok":true,"reward":0.8802},{"n":12,"raw":202.0,"best":234.8,"ok":true,"reward":0.827},{"n":13,"raw":236.9,"best":236.9,"ok":true,"reward":0.8834}],"final_speedup":236.856158,"vs_oracle":0.766863}}},{"task_id":"wro-w8a16-groupdequant-matmul-sol","short_name":"W8A16 grouped-asym fused dequant-GEMM Triton kernel","infra_name":"w8a16_group_dequant_matmul","summary":"This is a matmul subsystem multiplying group-quantized int8 weights by fp16 activations: along the K dimension, every group_size rows share one (scale, zero-point) pair. The existing w8a16_matmul is correct but slow; you must make the dequantization and the matmul fast (ideally fusing away the intermediate dense weights). The inputs are the activation a, the int8 qweight, the scales/zeros, and group_size; the output is a@dequant(W). Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 6.42047): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"C","medium_topic":"C1","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":6.42047,"traj":[{"n":1,"raw":11.89,"best":11.89,"ok":true,"reward":1.426},{"n":2,"raw":11.77,"best":11.89,"ok":true,"reward":1.416},{"n":3,"raw":12.27,"best":12.27,"ok":true,"reward":1.455},{"n":4,"raw":12.24,"best":12.27,"ok":true,"reward":1.453},{"n":5,"raw":12.35,"best":12.35,"ok":true,"reward":1.462},{"n":6,"raw":13.89,"best":12.35,"ok":false,"reward":0.0},{"n":7,"raw":16.09,"best":12.35,"ok":false,"reward":0.0},{"n":8,"raw":16.1,"best":12.35,"ok":false,"reward":0.0},{"n":9,"raw":14.56,"best":12.35,"ok":false,"reward":0.0},{"n":10,"raw":14.44,"best":12.35,"ok":false,"reward":0.0},{"n":11,"raw":13.53,"best":12.35,"ok":false,"reward":0.0},{"n":12,"raw":13.65,"best":12.35,"ok":false,"reward":0.0},{"n":13,"raw":12.19,"best":12.35,"ok":true,"reward":1.449},{"n":14,"raw":14.77,"best":12.35,"ok":false,"reward":0.0},{"n":15,"raw":16.03,"best":12.35,"ok":false,"reward":0.0}],"final_speedup":12.348113,"vs_oracle":1.923241},"kimi-k3":{"oracle":6.42047,"traj":[{"n":1,"raw":7.789,"best":7.789,"ok":true,"reward":1.107},{"n":2,"raw":10.63,"best":10.63,"ok":true,"reward":1.328},{"n":3,"raw":10.94,"best":10.94,"ok":true,"reward":1.352},{"n":4,"raw":10.91,"best":10.94,"ok":true,"reward":1.35},{"n":5,"raw":11.07,"best":11.07,"ok":true,"reward":1.362},{"n":6,"raw":11.7,"best":11.7,"ok":true,"reward":1.411},{"n":7,"raw":13.52,"best":11.7,"ok":false,"reward":0.0},{"n":8,"raw":13.11,"best":11.7,"ok":false,"reward":0.0},{"n":9,"raw":13.04,"best":11.7,"ok":false,"reward":0.0},{"n":10,"raw":11.84,"best":11.84,"ok":true,"reward":1.422},{"n":11,"raw":14.64,"best":11.84,"ok":false,"reward":0.0},{"n":12,"raw":11.23,"best":11.84,"ok":true,"reward":1.374},{"n":13,"raw":11.43,"best":11.84,"ok":true,"reward":1.39},{"n":14,"raw":11.52,"best":11.84,"ok":true,"reward":1.397},{"n":15,"raw":10.59,"best":11.84,"ok":true,"reward":1.325}],"final_speedup":11.618978,"vs_oracle":1.809677},"gpt-5.6":{"oracle":6.42047,"traj":[{"n":4,"raw":8.453,"best":8.453,"ok":true,"reward":1.158},{"n":13,"raw":8.435,"best":8.453,"ok":true,"reward":1.157},{"n":14,"raw":8.465,"best":8.465,"ok":true,"reward":1.159}],"final_speedup":8.463956,"vs_oracle":1.318277},"glm-5.2":{"oracle":6.42047,"traj":[{"n":1,"raw":3.043,"best":3.043,"ok":true,"reward":0.737},{"n":2,"raw":7.784,"best":7.784,"ok":true,"reward":1.106},{"n":3,"raw":7.804,"best":7.804,"ok":true,"reward":1.108},{"n":4,"raw":7.779,"best":7.804,"ok":true,"reward":1.106},{"n":5,"raw":7.766,"best":7.804,"ok":true,"reward":1.105},{"n":6,"raw":7.78,"best":7.804,"ok":true,"reward":1.106}],"final_speedup":7.776603,"vs_oracle":1.21122},"sonnet-5":{"oracle":6.42047,"traj":[{"n":1,"raw":4.776,"best":4.776,"ok":true,"reward":0.872},{"n":2,"raw":5.356,"best":5.356,"ok":true,"reward":0.9171},{"n":3,"raw":5.511,"best":5.511,"ok":true,"reward":0.9291},{"n":4,"raw":8.146,"best":8.146,"ok":true,"reward":1.134}],"final_speedup":8.14618,"vs_oracle":1.268782},"qwen3.7":{"oracle":6.42047,"traj":[{"n":1,"raw":4.678,"best":4.678,"ok":true,"reward":0.8643},{"n":2,"raw":3.51,"best":4.678,"ok":true,"reward":0.7733},{"n":3,"raw":5.651,"best":5.651,"ok":true,"reward":0.9401},{"n":4,"raw":5.733,"best":5.733,"ok":true,"reward":0.9464},{"n":5,"raw":4.37,"best":5.733,"ok":true,"reward":0.8403},{"n":6,"raw":4.607,"best":5.733,"ok":true,"reward":0.8588},{"n":7,"raw":7.397,"best":7.397,"ok":true,"reward":1.076},{"n":8,"raw":7.374,"best":7.397,"ok":true,"reward":1.074},{"n":9,"raw":6.014,"best":7.397,"ok":true,"reward":0.9684},{"n":10,"raw":6.79,"best":7.397,"ok":true,"reward":1.029},{"n":11,"raw":6.569,"best":7.397,"ok":true,"reward":1.012},{"n":12,"raw":6.581,"best":7.397,"ok":true,"reward":1.012},{"n":13,"raw":5.018,"best":7.397,"ok":true,"reward":0.8908},{"n":14,"raw":5.647,"best":7.397,"ok":true,"reward":0.9397},{"n":15,"raw":7.355,"best":7.397,"ok":true,"reward":1.073}],"final_speedup":7.359879,"vs_oracle":1.146315},"qwen3.8":{"oracle":6.42047,"traj":[{"n":1,"raw":7.798,"best":7.798,"ok":true,"reward":1.107},{"n":2,"raw":6.479,"best":7.798,"ok":true,"reward":1.005},{"n":3,"raw":6.459,"best":7.798,"ok":true,"reward":1.003},{"n":4,"raw":6.212,"best":7.798,"ok":true,"reward":0.9837},{"n":5,"raw":7.025,"best":7.798,"ok":true,"reward":1.047},{"n":6,"raw":7.806,"best":7.806,"ok":true,"reward":1.108},{"n":7,"raw":8.871,"best":8.871,"ok":true,"reward":1.191},{"n":8,"raw":8.951,"best":8.951,"ok":true,"reward":1.197},{"n":9,"raw":9.809,"best":9.809,"ok":true,"reward":1.264},{"n":10,"raw":9.774,"best":9.809,"ok":true,"reward":1.261}],"final_speedup":9.713506,"vs_oracle":1.512896},"deepseek":{"oracle":6.42047,"traj":[{"n":1,"raw":0.01762,"best":0.01762,"ok":true,"reward":0.5014},{"n":2,"raw":0.07109,"best":0.07109,"ok":true,"reward":0.5055},{"n":3,"raw":0.1237,"best":0.1237,"ok":true,"reward":0.5096},{"n":4,"raw":0.1548,"best":0.1548,"ok":true,"reward":0.5121},{"n":5,"raw":0.2633,"best":0.2633,"ok":true,"reward":0.5205},{"n":6,"raw":0.4872,"best":0.4872,"ok":true,"reward":0.5379},{"n":7,"raw":0.0,"best":0.4872,"ok":true,"reward":0.4643},{"n":8,"raw":0.07773,"best":0.4872,"ok":true,"reward":0.5061},{"n":9,"raw":0.07957,"best":0.4872,"ok":true,"reward":0.5062},{"n":10,"raw":0.0,"best":0.4872,"ok":true,"reward":0.3928},{"n":11,"raw":0.5585,"best":0.5585,"ok":true,"reward":0.5435},{"n":12,"raw":0.0,"best":0.5585,"ok":true,"reward":null},{"n":13,"raw":0.4535,"best":0.5585,"ok":true,"reward":0.5353},{"n":14,"raw":0.5799,"best":0.5799,"ok":true,"reward":0.5452},{"n":15,"raw":0.5826,"best":0.5826,"ok":true,"reward":0.5454},{"n":16,"raw":0.5666,"best":0.5826,"ok":true,"reward":0.5441}],"final_speedup":0.581323,"vs_oracle":0.090542}}},{"task_id":"wro-ssm-ssd-chunkscan","short_name":"Mamba-2 SSD chunked-scan kernel from scratch","infra_name":"mamba2_ssd_chunk_scan","summary":"This is vLLM's Mamba-2 SSD (state-space dual) sequence-scan subsystem, where five coupled files make up the full chunk-scan pipeline. The existing implementation is correct but slow; you must make it fast within a given tolerance, and your changes may only land within the declared scope. The inputs are the sequence's state-space parameters and input stream; the output is the scanned output (including the cross-chunk state passing). Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 467.4): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"D","medium_topic":"D5","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":467.4,"traj":[{"n":1,"raw":256.5,"best":256.5,"ok":true,"reward":256.5},{"n":2,"raw":339.1,"best":339.1,"ok":true,"reward":339.1},{"n":3,"raw":514.0,"best":514.0,"ok":true,"reward":514.0},{"n":4,"raw":609.8,"best":609.8,"ok":true,"reward":609.8},{"n":5,"raw":623.1,"best":623.1,"ok":true,"reward":623.1},{"n":6,"raw":711.9,"best":711.9,"ok":true,"reward":711.9},{"n":7,"raw":711.7,"best":711.9,"ok":true,"reward":711.7},{"n":8,"raw":1652.0,"best":1652.0,"ok":true,"reward":1652.0},{"n":9,"raw":1987.0,"best":1987.0,"ok":true,"reward":1987.0},{"n":10,"raw":2184.0,"best":2184.0,"ok":true,"reward":2184.0},{"n":11,"raw":2244.0,"best":2244.0,"ok":true,"reward":2244.0},{"n":12,"raw":2199.0,"best":2244.0,"ok":true,"reward":2199.0},{"n":13,"raw":2292.0,"best":2292.0,"ok":true,"reward":2292.0},{"n":14,"raw":2605.0,"best":2605.0,"ok":true,"reward":2605.0},{"n":15,"raw":2583.0,"best":2605.0,"ok":true,"reward":2583.0}],"final_speedup":2511.714226,"vs_oracle":5.3738},"kimi-k3":{"oracle":467.4,"traj":[{"n":1,"raw":108.3,"best":108.3,"ok":true,"reward":108.3},{"n":2,"raw":353.0,"best":353.0,"ok":true,"reward":353.0},{"n":3,"raw":410.3,"best":410.3,"ok":true,"reward":410.3},{"n":4,"raw":440.8,"best":440.8,"ok":true,"reward":440.8},{"n":5,"raw":882.8,"best":882.8,"ok":true,"reward":882.8},{"n":6,"raw":815.0,"best":882.8,"ok":true,"reward":815.0},{"n":7,"raw":783.4,"best":882.8,"ok":true,"reward":783.4},{"n":8,"raw":907.2,"best":907.2,"ok":true,"reward":907.2},{"n":9,"raw":983.5,"best":983.5,"ok":true,"reward":983.5},{"n":10,"raw":776.1,"best":983.5,"ok":true,"reward":776.1},{"n":11,"raw":840.5,"best":983.5,"ok":true,"reward":840.5},{"n":12,"raw":1182.0,"best":1182.0,"ok":true,"reward":1182.0},{"n":13,"raw":1282.0,"best":1282.0,"ok":true,"reward":1282.0},{"n":14,"raw":1136.0,"best":1282.0,"ok":true,"reward":1136.0}],"final_speedup":1158.589633,"vs_oracle":2.478797},"gpt-5.6":{"oracle":467.4,"traj":[{"n":3,"raw":147.3,"best":147.3,"ok":true,"reward":147.3},{"n":15,"raw":98.63,"best":147.3,"ok":true,"reward":98.63},{"n":16,"raw":157.5,"best":157.5,"ok":true,"reward":157.5}],"final_speedup":144.408554,"vs_oracle":0.308961},"glm-5.2":{"oracle":467.4,"traj":[{"n":1,"raw":169.5,"best":169.5,"ok":true,"reward":169.5},{"n":2,"raw":182.8,"best":182.8,"ok":true,"reward":182.8},{"n":3,"raw":183.9,"best":183.9,"ok":true,"reward":183.9},{"n":4,"raw":213.2,"best":213.2,"ok":true,"reward":213.2},{"n":5,"raw":216.6,"best":216.6,"ok":true,"reward":216.6},{"n":6,"raw":315.0,"best":315.0,"ok":true,"reward":315.0},{"n":7,"raw":318.3,"best":318.3,"ok":true,"reward":318.3},{"n":8,"raw":329.0,"best":329.0,"ok":true,"reward":329.0},{"n":9,"raw":366.9,"best":366.9,"ok":true,"reward":366.9},{"n":10,"raw":374.4,"best":374.4,"ok":true,"reward":374.4},{"n":11,"raw":387.9,"best":387.9,"ok":true,"reward":387.9},{"n":12,"raw":348.2,"best":387.9,"ok":true,"reward":348.2}],"final_speedup":391.003986,"vs_oracle":0.836551},"sonnet-5":{"oracle":467.4,"traj":[{"n":1,"raw":51.23,"best":51.23,"ok":true,"reward":51.23},{"n":2,"raw":103.2,"best":103.2,"ok":true,"reward":103.2},{"n":3,"raw":99.82,"best":103.2,"ok":true,"reward":99.82},{"n":4,"raw":133.5,"best":133.5,"ok":true,"reward":133.5},{"n":5,"raw":152.7,"best":152.7,"ok":true,"reward":152.7},{"n":6,"raw":148.0,"best":152.7,"ok":true,"reward":148.0},{"n":7,"raw":152.0,"best":152.7,"ok":true,"reward":152.0},{"n":8,"raw":161.0,"best":161.0,"ok":true,"reward":161.0},{"n":9,"raw":161.2,"best":161.2,"ok":true,"reward":161.2},{"n":10,"raw":175.5,"best":175.5,"ok":true,"reward":175.5},{"n":11,"raw":166.2,"best":175.5,"ok":true,"reward":166.2},{"n":12,"raw":163.4,"best":175.5,"ok":true,"reward":163.4}],"final_speedup":165.137551,"vs_oracle":0.353311},"qwen3.7":{"oracle":467.4,"traj":[{"n":1,"raw":78.61,"best":78.61,"ok":true,"reward":78.61},{"n":2,"raw":84.43,"best":84.43,"ok":true,"reward":84.43},{"n":3,"raw":71.04,"best":84.43,"ok":true,"reward":71.04},{"n":4,"raw":90.3,"best":90.3,"ok":true,"reward":90.3},{"n":5,"raw":181.7,"best":181.7,"ok":true,"reward":181.7},{"n":6,"raw":174.7,"best":181.7,"ok":true,"reward":174.7},{"n":7,"raw":187.8,"best":187.8,"ok":true,"reward":187.8},{"n":8,"raw":177.7,"best":187.8,"ok":true,"reward":177.7},{"n":9,"raw":176.9,"best":187.8,"ok":true,"reward":176.9},{"n":10,"raw":173.2,"best":187.8,"ok":true,"reward":173.2},{"n":11,"raw":172.9,"best":187.8,"ok":true,"reward":172.9},{"n":12,"raw":174.4,"best":187.8,"ok":true,"reward":174.4},{"n":13,"raw":173.2,"best":187.8,"ok":true,"reward":173.2},{"n":14,"raw":175.8,"best":187.8,"ok":true,"reward":175.8},{"n":15,"raw":173.4,"best":187.8,"ok":true,"reward":173.4}],"final_speedup":174.642166,"vs_oracle":0.373646},"qwen3.8":{"oracle":467.4,"traj":[{"n":1,"raw":766.2,"best":766.2,"ok":true,"reward":766.2},{"n":2,"raw":903.2,"best":903.2,"ok":true,"reward":903.2},{"n":3,"raw":798.6,"best":903.2,"ok":true,"reward":798.6},{"n":4,"raw":909.9,"best":909.9,"ok":true,"reward":909.9},{"n":5,"raw":891.4,"best":909.9,"ok":true,"reward":891.4},{"n":6,"raw":0.0,"best":909.9,"ok":false,"reward":0.0},{"n":7,"raw":866.8,"best":909.9,"ok":true,"reward":866.8},{"n":8,"raw":889.9,"best":909.9,"ok":true,"reward":889.9},{"n":9,"raw":921.3,"best":921.3,"ok":true,"reward":921.3},{"n":10,"raw":916.2,"best":921.3,"ok":true,"reward":916.2},{"n":11,"raw":950.1,"best":950.1,"ok":true,"reward":950.1},{"n":12,"raw":889.3,"best":950.1,"ok":true,"reward":889.3},{"n":13,"raw":428.9,"best":950.1,"ok":true,"reward":428.9},{"n":14,"raw":242.9,"best":950.1,"ok":true,"reward":242.9},{"n":15,"raw":886.4,"best":950.1,"ok":true,"reward":886.4}],"final_speedup":887.784442,"vs_oracle":null},"deepseek":{"oracle":467.4,"traj":[{"n":1,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":2,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":3,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":4,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":5,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":6,"raw":1.142,"best":1.142,"ok":true,"reward":1.142},{"n":7,"raw":0.0,"best":1.142,"ok":false,"reward":0.0},{"n":8,"raw":0.0,"best":1.142,"ok":false,"reward":0.0},{"n":9,"raw":0.0,"best":1.142,"ok":false,"reward":0.0},{"n":10,"raw":1.228,"best":1.228,"ok":true,"reward":1.228},{"n":11,"raw":0.0,"best":1.228,"ok":false,"reward":0.0},{"n":12,"raw":0.0,"best":1.228,"ok":false,"reward":0.0}],"final_speedup":1.094406,"vs_oracle":0.002341}}},{"task_id":"wro-sp24mm-matmul-sol","short_name":"2:4 semi-structured sparse GEMM Triton kernel","infra_name":"sparse_2_4_fp16_matmul","summary":"This is a matmul subsystem multiplying 2:4 semi-structured sparse weights by fp16 activations: within every 4 consecutive K rows exactly 2 are nonzero, and the weights are stored in a compressed format (values + 2-bit position metadata). The existing sp24mm_matmul is correct but slow; you must make it fast, and you may not densify the weights. The inputs are the activation a, the compressed weight values w_vals, and the metadata w_meta; the output is a@W. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 10.2846): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"C","medium_topic":"C2","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":10.28458,"traj":[{"n":1,"raw":4.832,"best":4.832,"ok":true,"reward":0.7349},{"n":2,"raw":4.859,"best":4.859,"ok":true,"reward":0.7362},{"n":3,"raw":13.31,"best":13.31,"ok":true,"reward":1.147},{"n":4,"raw":15.41,"best":15.41,"ok":true,"reward":1.249},{"n":5,"raw":16.93,"best":16.93,"ok":true,"reward":1.323},{"n":6,"raw":17.44,"best":17.44,"ok":true,"reward":1.348},{"n":7,"raw":17.43,"best":17.44,"ok":true,"reward":1.348},{"n":8,"raw":17.36,"best":17.44,"ok":true,"reward":1.344}],"final_speedup":17.442733,"vs_oracle":1.696008},"kimi-k3":{"oracle":10.28458,"traj":[{"n":1,"raw":3.203,"best":3.203,"ok":true,"reward":0.6557},{"n":2,"raw":7.967,"best":7.967,"ok":true,"reward":0.8873},{"n":3,"raw":7.301,"best":7.967,"ok":true,"reward":0.855},{"n":4,"raw":7.309,"best":7.967,"ok":true,"reward":0.8553},{"n":5,"raw":7.272,"best":7.967,"ok":true,"reward":0.8536},{"n":6,"raw":7.947,"best":7.967,"ok":true,"reward":0.8863}],"final_speedup":7.925026,"vs_oracle":0.770574},"gpt-5.6":{"oracle":10.28458,"traj":[{"n":6,"raw":3.752,"best":3.752,"ok":true,"reward":0.6824},{"n":14,"raw":3.651,"best":3.752,"ok":true,"reward":0.6775}],"final_speedup":3.737377,"vs_oracle":0.363396},"glm-5.2":{"oracle":10.28458,"traj":[{"n":1,"raw":3.529,"best":3.529,"ok":true,"reward":0.6716},{"n":2,"raw":5.587,"best":5.587,"ok":true,"reward":0.7716},{"n":3,"raw":12.32,"best":12.32,"ok":true,"reward":1.099},{"n":4,"raw":12.42,"best":12.42,"ok":true,"reward":1.104}],"final_speedup":12.532507,"vs_oracle":1.218573},"sonnet-5":{"oracle":10.28458,"traj":[{"n":1,"raw":11.19,"best":11.19,"ok":true,"reward":1.044},{"n":2,"raw":5.039,"best":11.19,"ok":true,"reward":0.745},{"n":3,"raw":11.16,"best":11.19,"ok":true,"reward":1.042},{"n":4,"raw":5.547,"best":11.19,"ok":true,"reward":0.7697},{"n":5,"raw":12.6,"best":12.6,"ok":true,"reward":1.113},{"n":6,"raw":13.36,"best":13.36,"ok":true,"reward":1.149},{"n":7,"raw":5.353,"best":13.36,"ok":true,"reward":0.7602},{"n":8,"raw":6.673,"best":13.36,"ok":true,"reward":0.8244},{"n":9,"raw":6.536,"best":13.36,"ok":true,"reward":0.8178},{"n":10,"raw":8.299,"best":13.36,"ok":true,"reward":0.9035},{"n":11,"raw":13.47,"best":13.47,"ok":true,"reward":1.155},{"n":12,"raw":12.52,"best":13.47,"ok":true,"reward":1.109},{"n":13,"raw":13.52,"best":13.52,"ok":true,"reward":1.157}],"final_speedup":13.399065,"vs_oracle":1.302831},"qwen3.7":{"oracle":10.28458,"traj":[{"n":1,"raw":0.1024,"best":0.1024,"ok":true,"reward":0.505},{"n":2,"raw":0.2416,"best":0.2416,"ok":true,"reward":0.5117},{"n":3,"raw":0.0,"best":0.2416,"ok":true,"reward":0.0},{"n":4,"raw":0.2178,"best":0.2416,"ok":true,"reward":0.5106},{"n":5,"raw":0.2131,"best":0.2416,"ok":true,"reward":0.5104},{"n":6,"raw":5.739,"best":5.739,"ok":true,"reward":0.779},{"n":7,"raw":6.395,"best":6.395,"ok":true,"reward":0.8109},{"n":8,"raw":7.064,"best":7.064,"ok":true,"reward":0.8434},{"n":9,"raw":7.144,"best":7.144,"ok":true,"reward":0.8473},{"n":10,"raw":0.0,"best":7.144,"ok":true,"reward":0.0},{"n":11,"raw":5.511,"best":7.144,"ok":true,"reward":0.7679},{"n":12,"raw":0.0,"best":7.144,"ok":true,"reward":0.0},{"n":13,"raw":7.165,"best":7.165,"ok":true,"reward":0.8484},{"n":14,"raw":7.704,"best":7.704,"ok":true,"reward":0.8745},{"n":15,"raw":3.887,"best":7.704,"ok":true,"reward":0.689},{"n":16,"raw":7.286,"best":7.704,"ok":true,"reward":0.8542}],"final_speedup":7.681216,"vs_oracle":0.746867},"qwen3.8":{"oracle":10.28458,"traj":[{"n":1,"raw":10.64,"best":10.64,"ok":true,"reward":1.017},{"n":2,"raw":11.67,"best":11.67,"ok":true,"reward":1.067},{"n":3,"raw":11.66,"best":11.67,"ok":true,"reward":1.067},{"n":4,"raw":13.63,"best":13.63,"ok":true,"reward":1.162}],"final_speedup":13.539552,"vs_oracle":1.316491},"deepseek":{"oracle":10.28458,"traj":[{"n":1,"raw":0.06916,"best":0.06916,"ok":true,"reward":0.5034},{"n":2,"raw":0.5245,"best":0.5245,"ok":true,"reward":0.5255},{"n":3,"raw":0.0,"best":0.5245,"ok":true,"reward":0.3572},{"n":4,"raw":0.0,"best":0.5245,"ok":true,"reward":0.3572},{"n":5,"raw":0.7312,"best":0.7312,"ok":true,"reward":0.5355},{"n":6,"raw":0.6441,"best":0.7312,"ok":true,"reward":0.5313},{"n":7,"raw":0.7444,"best":0.7444,"ok":true,"reward":0.5362},{"n":8,"raw":0.5357,"best":0.7444,"ok":true,"reward":0.526},{"n":9,"raw":0.7042,"best":0.7444,"ok":true,"reward":0.5342},{"n":10,"raw":0.7224,"best":0.7444,"ok":true,"reward":0.5351},{"n":11,"raw":0.7494,"best":0.7494,"ok":true,"reward":0.5364},{"n":12,"raw":0.6227,"best":0.7494,"ok":true,"reward":0.5303},{"n":13,"raw":0.0,"best":0.7494,"ok":true,"reward":0.4643},{"n":14,"raw":0.0,"best":0.7494,"ok":true,"reward":0.4643},{"n":15,"raw":0.747,"best":0.7494,"ok":true,"reward":0.5363},{"n":16,"raw":0.7463,"best":0.7494,"ok":true,"reward":0.5363}],"final_speedup":0.744614,"vs_oracle":0.072401}}},{"task_id":"wro-lora-punica-fused","short_name":"Punica multi-LoRA segmented fused GEMM kernel","infra_name":"punica_multilora_shrink_expand","summary":"This is vLLM's Punica-style batched multi-LoRA application subsystem (a two-stage shrink → expand path), spanning three Triton kernel files. The existing implementation is correct but slow; you must make it fast within a 2% relative tolerance (fp16 inputs, fp32 accumulation). The inputs are the batched activations together with the A/B weights of the several LoRAs and their grouping information; the output is the result with each LoRA's own contribution added on. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 16.59): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"B","medium_topic":"B9","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":16.59,"traj":[{"n":1,"raw":18.37,"best":18.37,"ok":true,"reward":18.37},{"n":2,"raw":18.73,"best":18.73,"ok":true,"reward":18.73},{"n":3,"raw":36.05,"best":36.05,"ok":true,"reward":36.05},{"n":4,"raw":35.49,"best":36.05,"ok":true,"reward":35.49},{"n":5,"raw":36.22,"best":36.22,"ok":true,"reward":36.22},{"n":6,"raw":35.5,"best":36.22,"ok":true,"reward":35.5},{"n":7,"raw":37.79,"best":37.79,"ok":true,"reward":37.79},{"n":8,"raw":36.41,"best":37.79,"ok":true,"reward":36.41},{"n":9,"raw":36.47,"best":37.79,"ok":true,"reward":36.47},{"n":10,"raw":36.75,"best":37.79,"ok":true,"reward":36.75},{"n":11,"raw":36.09,"best":37.79,"ok":true,"reward":36.09},{"n":12,"raw":36.03,"best":37.79,"ok":true,"reward":36.03},{"n":13,"raw":36.35,"best":37.79,"ok":true,"reward":36.35}],"final_speedup":35.453852,"vs_oracle":2.137062},"kimi-k3":{"oracle":16.59,"traj":[{"n":1,"raw":2.419,"best":2.419,"ok":true,"reward":2.419},{"n":2,"raw":3.913,"best":3.913,"ok":true,"reward":3.913},{"n":3,"raw":0.0,"best":3.913,"ok":false,"reward":0.0},{"n":4,"raw":17.48,"best":17.48,"ok":true,"reward":17.48},{"n":5,"raw":21.49,"best":21.49,"ok":true,"reward":21.49},{"n":6,"raw":21.18,"best":21.49,"ok":true,"reward":21.18},{"n":7,"raw":29.55,"best":29.55,"ok":true,"reward":29.55},{"n":8,"raw":36.24,"best":36.24,"ok":true,"reward":36.24},{"n":9,"raw":37.47,"best":37.47,"ok":true,"reward":37.47},{"n":10,"raw":36.05,"best":37.47,"ok":true,"reward":36.05},{"n":11,"raw":36.06,"best":37.47,"ok":true,"reward":36.06},{"n":12,"raw":36.65,"best":37.47,"ok":true,"reward":36.65},{"n":13,"raw":36.77,"best":37.47,"ok":true,"reward":36.77}],"final_speedup":36.509442,"vs_oracle":2.20069},"gpt-5.6":{"oracle":16.59,"traj":[{"n":5,"raw":22.49,"best":22.49,"ok":true,"reward":22.49},{"n":14,"raw":18.06,"best":22.49,"ok":true,"reward":18.06}],"final_speedup":21.67669,"vs_oracle":1.306612},"glm-5.2":{"oracle":16.59,"traj":[{"n":1,"raw":1.874,"best":1.874,"ok":true,"reward":1.874},{"n":2,"raw":8.387,"best":8.387,"ok":true,"reward":8.387},{"n":3,"raw":21.54,"best":21.54,"ok":true,"reward":21.54},{"n":4,"raw":22.23,"best":22.23,"ok":true,"reward":22.23},{"n":5,"raw":22.14,"best":22.23,"ok":true,"reward":22.14},{"n":6,"raw":22.36,"best":22.36,"ok":true,"reward":22.36},{"n":7,"raw":24.93,"best":24.93,"ok":true,"reward":24.93},{"n":8,"raw":25.02,"best":25.02,"ok":true,"reward":25.02}],"final_speedup":24.958024,"vs_oracle":1.504402},"sonnet-5":{"oracle":16.59,"traj":[{"n":1,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":2,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":3,"raw":4.935,"best":4.935,"ok":true,"reward":4.935},{"n":4,"raw":5.311,"best":5.311,"ok":true,"reward":5.311},{"n":5,"raw":5.348,"best":5.348,"ok":true,"reward":5.348},{"n":6,"raw":4.982,"best":5.348,"ok":true,"reward":4.982},{"n":7,"raw":5.483,"best":5.483,"ok":true,"reward":5.483},{"n":8,"raw":9.75,"best":9.75,"ok":true,"reward":9.75},{"n":9,"raw":9.328,"best":9.75,"ok":true,"reward":9.328},{"n":10,"raw":10.1,"best":10.1,"ok":true,"reward":10.1},{"n":11,"raw":10.3,"best":10.3,"ok":true,"reward":10.3},{"n":12,"raw":9.24,"best":10.3,"ok":true,"reward":9.24},{"n":13,"raw":7.34,"best":10.3,"ok":true,"reward":7.34},{"n":14,"raw":10.27,"best":10.3,"ok":true,"reward":10.27}],"final_speedup":10.282851,"vs_oracle":0.619822},"qwen3.7":{"oracle":16.59,"traj":[{"n":1,"raw":1.017,"best":1.017,"ok":true,"reward":1.017},{"n":2,"raw":0.8807,"best":1.017,"ok":true,"reward":0.8807},{"n":3,"raw":1.157,"best":1.157,"ok":true,"reward":1.157},{"n":4,"raw":0.9463,"best":1.157,"ok":true,"reward":0.9463},{"n":5,"raw":1.885,"best":1.885,"ok":true,"reward":1.885},{"n":6,"raw":1.973,"best":1.973,"ok":true,"reward":1.973},{"n":7,"raw":1.794,"best":1.973,"ok":true,"reward":1.794},{"n":8,"raw":1.927,"best":1.973,"ok":true,"reward":1.927},{"n":9,"raw":1.898,"best":1.973,"ok":true,"reward":1.898},{"n":10,"raw":2.124,"best":2.124,"ok":true,"reward":2.124},{"n":11,"raw":1.929,"best":2.124,"ok":true,"reward":1.929},{"n":12,"raw":1.962,"best":2.124,"ok":true,"reward":1.962},{"n":13,"raw":1.947,"best":2.124,"ok":true,"reward":1.947},{"n":14,"raw":1.875,"best":2.124,"ok":true,"reward":1.875}],"final_speedup":1.882253,"vs_oracle":0.113457},"qwen3.8":{"oracle":16.59,"traj":[{"n":1,"raw":13.67,"best":13.67,"ok":true,"reward":13.67},{"n":2,"raw":19.64,"best":19.64,"ok":true,"reward":19.64},{"n":3,"raw":19.64,"best":19.64,"ok":true,"reward":19.64},{"n":4,"raw":17.56,"best":19.64,"ok":true,"reward":17.56},{"n":5,"raw":16.75,"best":19.64,"ok":true,"reward":16.75},{"n":6,"raw":19.9,"best":19.9,"ok":true,"reward":19.9},{"n":7,"raw":19.26,"best":19.9,"ok":true,"reward":19.26},{"n":8,"raw":20.43,"best":20.43,"ok":true,"reward":20.43},{"n":9,"raw":18.08,"best":20.43,"ok":true,"reward":18.08},{"n":10,"raw":23.52,"best":23.52,"ok":true,"reward":23.52},{"n":11,"raw":25.33,"best":25.33,"ok":true,"reward":25.33},{"n":12,"raw":27.86,"best":27.86,"ok":true,"reward":27.86},{"n":13,"raw":12.35,"best":27.86,"ok":true,"reward":12.35},{"n":14,"raw":27.05,"best":27.86,"ok":true,"reward":27.05},{"n":15,"raw":27.11,"best":27.86,"ok":true,"reward":27.11},{"n":16,"raw":27.85,"best":27.86,"ok":true,"reward":27.85}],"final_speedup":27.207341,"vs_oracle":null},"deepseek":{"oracle":16.59,"traj":[{"n":1,"raw":1.315,"best":1.315,"ok":true,"reward":1.315},{"n":2,"raw":0.0,"best":1.315,"ok":false,"reward":0.0},{"n":3,"raw":1.481,"best":1.481,"ok":true,"reward":1.481},{"n":4,"raw":3.149,"best":3.149,"ok":true,"reward":3.149},{"n":5,"raw":4.088,"best":4.088,"ok":true,"reward":4.088},{"n":6,"raw":6.434,"best":6.434,"ok":true,"reward":6.434},{"n":7,"raw":0.0,"best":6.434,"ok":false,"reward":0.0},{"n":8,"raw":3.031,"best":6.434,"ok":true,"reward":3.031},{"n":9,"raw":2.647,"best":6.434,"ok":true,"reward":2.647},{"n":10,"raw":0.0,"best":6.434,"ok":false,"reward":0.0},{"n":11,"raw":0.0,"best":6.434,"ok":false,"reward":0.0},{"n":12,"raw":6.316,"best":6.434,"ok":true,"reward":6.316},{"n":13,"raw":0.0,"best":6.434,"ok":false,"reward":0.0},{"n":14,"raw":6.429,"best":6.434,"ok":true,"reward":6.429},{"n":15,"raw":5.761,"best":6.434,"ok":true,"reward":5.761}],"final_speedup":6.333061,"vs_oracle":0.38174}}},{"task_id":"wro-fla-abc-chunkscan-sol","short_name":"ABC (Bounded-memory Attention) chunked-scan kernel","infra_name":"abc_gated_slot_recurrence","summary":"This is chunk_abc, a two-stage gated linear-recurrence operator in flash-linear-attention with bounded slot memory. The existing implementation is correct but slow; you must speed up this chunk scan while preserving its numerical behavior. The input is per-step q/k/v and slot logits, and the output is the output stream (exposed via torch.autograd.Function). Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 578.38): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"D","medium_topic":"D5","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":578.380265,"traj":[{"n":1,"raw":611.1,"best":611.1,"ok":true,"reward":1.028},{"n":2,"raw":949.6,"best":949.6,"ok":true,"reward":1.321},{"n":3,"raw":1151.0,"best":1151.0,"ok":true,"reward":1.495},{"n":4,"raw":1122.0,"best":1151.0,"ok":true,"reward":1.47},{"n":5,"raw":1201.0,"best":1151.0,"ok":false,"reward":0.0},{"n":6,"raw":1560.0,"best":1151.0,"ok":false,"reward":0.0},{"n":7,"raw":1529.0,"best":1151.0,"ok":false,"reward":0.0},{"n":8,"raw":1467.0,"best":1151.0,"ok":false,"reward":0.0},{"n":9,"raw":1577.0,"best":1151.0,"ok":false,"reward":0.0},{"n":10,"raw":1051.0,"best":1151.0,"ok":true,"reward":1.409},{"n":11,"raw":1266.0,"best":1151.0,"ok":false,"reward":0.0},{"n":12,"raw":1420.0,"best":1151.0,"ok":false,"reward":0.0},{"n":13,"raw":1385.0,"best":1151.0,"ok":false,"reward":0.0}],"final_speedup":1176.07052,"vs_oracle":2.033386},"kimi-k3":{"oracle":578.380265,"traj":[{"n":1,"raw":136.5,"best":136.5,"ok":true,"reward":0.618}],"final_speedup":1053.841082,"vs_oracle":1.822056},"gpt-5.6":{"oracle":578.380265,"traj":[{"n":11,"raw":131.1,"best":131.1,"ok":true,"reward":0.6134},{"n":15,"raw":114.1,"best":131.1,"ok":true,"reward":0.5986},{"n":16,"raw":125.8,"best":131.1,"ok":true,"reward":0.6088}],"final_speedup":124.738685,"vs_oracle":0.215669},"glm-5.2":{"oracle":578.380265,"traj":[{"n":1,"raw":33.52,"best":33.52,"ok":true,"reward":0.529},{"n":2,"raw":150.6,"best":150.6,"ok":true,"reward":0.6302},{"n":3,"raw":0.0,"best":150.6,"ok":true,"reward":0.0},{"n":4,"raw":278.4,"best":278.4,"ok":true,"reward":0.7407},{"n":5,"raw":344.0,"best":344.0,"ok":true,"reward":0.7974},{"n":6,"raw":360.1,"best":360.1,"ok":true,"reward":0.8113}],"final_speedup":341.594798,"vs_oracle":0.590606},"sonnet-5":{"oracle":578.380265,"traj":[{"n":1,"raw":28.61,"best":28.61,"ok":true,"reward":0.5247},{"n":2,"raw":57.49,"best":57.49,"ok":true,"reward":0.5497},{"n":3,"raw":109.5,"best":109.5,"ok":true,"reward":0.5947},{"n":4,"raw":106.5,"best":109.5,"ok":true,"reward":0.592},{"n":5,"raw":106.0,"best":109.5,"ok":true,"reward":0.5916},{"n":6,"raw":80.15,"best":109.5,"ok":true,"reward":0.5693},{"n":7,"raw":118.1,"best":118.1,"ok":true,"reward":0.6021},{"n":8,"raw":128.8,"best":128.8,"ok":true,"reward":0.6114},{"n":9,"raw":130.9,"best":130.9,"ok":true,"reward":0.6132},{"n":10,"raw":115.4,"best":130.9,"ok":true,"reward":0.5998},{"n":11,"raw":120.1,"best":130.9,"ok":true,"reward":0.6038},{"n":12,"raw":126.2,"best":130.9,"ok":true,"reward":0.6091},{"n":13,"raw":123.6,"best":130.9,"ok":true,"reward":0.6068}],"final_speedup":124.941582,"vs_oracle":0.21602},"qwen3.7":{"oracle":578.380265,"traj":[{"n":1,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":2,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":3,"raw":1.041,"best":1.041,"ok":true,"reward":0.5009},{"n":4,"raw":1.083,"best":1.083,"ok":true,"reward":0.5009},{"n":5,"raw":1.029,"best":1.083,"ok":true,"reward":0.5009},{"n":6,"raw":1.131,"best":1.131,"ok":true,"reward":0.501},{"n":7,"raw":1.179,"best":1.179,"ok":true,"reward":0.501},{"n":8,"raw":1.434,"best":1.434,"ok":true,"reward":0.5012},{"n":9,"raw":1.772,"best":1.772,"ok":true,"reward":0.5015},{"n":10,"raw":1.743,"best":1.772,"ok":true,"reward":0.5015},{"n":11,"raw":1.422,"best":1.772,"ok":true,"reward":0.5012},{"n":12,"raw":1.763,"best":1.772,"ok":true,"reward":0.5015},{"n":13,"raw":1.792,"best":1.792,"ok":true,"reward":0.5015},{"n":14,"raw":1.749,"best":1.792,"ok":true,"reward":0.5015}],"final_speedup":1.726733,"vs_oracle":0.002985},"qwen3.8":{"oracle":578.380265,"traj":[{"n":1,"raw":241.7,"best":241.7,"ok":true,"reward":0.7089},{"n":2,"raw":373.6,"best":373.6,"ok":true,"reward":0.823},{"n":3,"raw":387.1,"best":387.1,"ok":true,"reward":0.8347},{"n":4,"raw":513.9,"best":513.9,"ok":true,"reward":0.9442},{"n":5,"raw":754.4,"best":754.4,"ok":true,"reward":1.152},{"n":6,"raw":1208.0,"best":754.4,"ok":false,"reward":0.0},{"n":7,"raw":1237.0,"best":754.4,"ok":false,"reward":0.0},{"n":8,"raw":1239.0,"best":754.4,"ok":false,"reward":0.0},{"n":9,"raw":778.4,"best":778.4,"ok":true,"reward":1.173},{"n":10,"raw":736.6,"best":778.4,"ok":true,"reward":1.137},{"n":11,"raw":1168.0,"best":778.4,"ok":false,"reward":0.0},{"n":12,"raw":765.7,"best":778.4,"ok":true,"reward":1.162},{"n":13,"raw":1187.0,"best":778.4,"ok":false,"reward":0.0},{"n":14,"raw":825.1,"best":825.1,"ok":true,"reward":1.213},{"n":15,"raw":1430.0,"best":825.1,"ok":false,"reward":0.0}],"final_speedup":842.565793,"vs_oracle":1.456768},"deepseek":{"oracle":578.380265,"traj":[{"n":1,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":2,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":3,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":4,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":5,"raw":1.053,"best":1.053,"ok":true,"reward":0.5009},{"n":6,"raw":0.0,"best":1.053,"ok":true,"reward":0.0},{"n":7,"raw":1.119,"best":1.119,"ok":true,"reward":0.501},{"n":8,"raw":0.0,"best":1.119,"ok":true,"reward":0.0},{"n":9,"raw":0.0,"best":1.119,"ok":true,"reward":0.0},{"n":10,"raw":0.0,"best":1.119,"ok":true,"reward":0.0},{"n":11,"raw":1.139,"best":1.139,"ok":true,"reward":0.501},{"n":12,"raw":1.235,"best":1.235,"ok":true,"reward":0.5011},{"n":13,"raw":1.141,"best":1.235,"ok":true,"reward":0.501},{"n":14,"raw":0.0,"best":1.235,"ok":true,"reward":0.0},{"n":15,"raw":1.292,"best":1.292,"ok":true,"reward":0.5011},{"n":16,"raw":1.26,"best":1.292,"ok":true,"reward":0.5011}],"final_speedup":1.284169,"vs_oracle":0.00222}}},{"task_id":"wro-fla-gated-delta-chunkscan-sol","short_name":"Gated DeltaNet chunked WY/UT-transform kernel","infra_name":"gated_deltanet_chunk_scan","summary":"This is the chunk forward of Gated DeltaNet (gated delta-rule linear attention) in flash-linear-attention, spanning three coupled files. The existing implementation is correct but slow; you must make it fast: at each step the state first decays by a scalar log-domain forget gate, then undergoes a beta-weighted delta update. The input is q/k/v, the gate g, and beta, and the output is the output of chunk_gated_delta_rule (including the final state when needed). Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 822.62): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"D","medium_topic":"D5","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":822.619615,"traj":[{"n":1,"raw":486.5,"best":486.5,"ok":true,"reward":0.7957},{"n":2,"raw":766.7,"best":766.7,"ok":true,"reward":0.966},{"n":3,"raw":717.7,"best":766.7,"ok":true,"reward":0.9362},{"n":4,"raw":716.7,"best":766.7,"ok":true,"reward":0.9356},{"n":5,"raw":865.0,"best":865.0,"ok":true,"reward":1.026},{"n":6,"raw":962.2,"best":962.2,"ok":true,"reward":1.085},{"n":7,"raw":907.9,"best":962.2,"ok":true,"reward":1.052},{"n":8,"raw":887.8,"best":962.2,"ok":true,"reward":1.04},{"n":9,"raw":952.2,"best":962.2,"ok":true,"reward":1.079}],"final_speedup":876.643049,"vs_oracle":1.065672},"kimi-k3":{"oracle":822.619615,"traj":[{"n":1,"raw":747.8,"best":747.8,"ok":true,"reward":0.9545},{"n":2,"raw":842.2,"best":842.2,"ok":true,"reward":1.012},{"n":3,"raw":848.3,"best":848.3,"ok":true,"reward":1.016},{"n":4,"raw":863.9,"best":863.9,"ok":true,"reward":1.025},{"n":5,"raw":910.4,"best":910.4,"ok":true,"reward":1.053},{"n":6,"raw":905.7,"best":910.4,"ok":true,"reward":1.05},{"n":7,"raw":912.2,"best":912.2,"ok":true,"reward":1.054},{"n":8,"raw":884.4,"best":912.2,"ok":true,"reward":1.038},{"n":9,"raw":957.8,"best":957.8,"ok":true,"reward":1.082},{"n":10,"raw":899.0,"best":957.8,"ok":true,"reward":1.046},{"n":11,"raw":887.7,"best":957.8,"ok":true,"reward":1.04},{"n":12,"raw":894.8,"best":957.8,"ok":true,"reward":1.044}],"final_speedup":916.317494,"vs_oracle":1.113902},"gpt-5.6":{"oracle":822.619615,"traj":[{"n":9,"raw":89.2,"best":89.2,"ok":true,"reward":0.5542},{"n":15,"raw":87.21,"best":89.2,"ok":true,"reward":0.553}],"final_speedup":86.793608,"vs_oracle":0.105509},"glm-5.2":{"oracle":822.619615,"traj":[{"n":1,"raw":314.5,"best":314.5,"ok":true,"reward":0.6912},{"n":2,"raw":658.7,"best":658.7,"ok":true,"reward":0.9004},{"n":3,"raw":680.9,"best":680.9,"ok":true,"reward":0.9139},{"n":4,"raw":738.4,"best":738.4,"ok":true,"reward":0.9488},{"n":5,"raw":809.0,"best":809.0,"ok":true,"reward":0.9917},{"n":6,"raw":797.6,"best":809.0,"ok":true,"reward":0.9848}],"final_speedup":767.557265,"vs_oracle":0.933065},"sonnet-5":{"oracle":822.619615,"traj":[{"n":1,"raw":157.4,"best":157.4,"ok":true,"reward":0.5956},{"n":2,"raw":187.5,"best":187.5,"ok":true,"reward":0.614},{"n":3,"raw":205.0,"best":205.0,"ok":true,"reward":0.6246},{"n":4,"raw":203.0,"best":205.0,"ok":true,"reward":0.6234},{"n":5,"raw":221.2,"best":221.2,"ok":true,"reward":0.6344},{"n":6,"raw":257.7,"best":257.7,"ok":true,"reward":0.6567},{"n":7,"raw":267.8,"best":267.8,"ok":true,"reward":0.6628}],"final_speedup":271.751445,"vs_oracle":0.330349},"qwen3.7":{"oracle":822.619615,"traj":[{"n":1,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":2,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":3,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":4,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":5,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":6,"raw":0.0,"best":0.0,"ok":true,"reward":0.2917},{"n":7,"raw":72.81,"best":72.81,"ok":true,"reward":0.5443},{"n":8,"raw":81.05,"best":81.05,"ok":true,"reward":0.5493},{"n":9,"raw":66.71,"best":81.05,"ok":true,"reward":0.5405},{"n":10,"raw":41.37,"best":81.05,"ok":true,"reward":0.5251},{"n":11,"raw":81.48,"best":81.48,"ok":true,"reward":0.5495},{"n":12,"raw":82.66,"best":82.66,"ok":true,"reward":0.5502},{"n":13,"raw":82.76,"best":82.76,"ok":true,"reward":0.5503},{"n":14,"raw":0.0,"best":82.76,"ok":true,"reward":0.0},{"n":15,"raw":77.8,"best":82.76,"ok":true,"reward":0.5473},{"n":16,"raw":78.04,"best":82.76,"ok":true,"reward":0.5474}],"final_speedup":81.782302,"vs_oracle":0.099417},"qwen3.8":{"oracle":822.619615,"traj":[{"n":1,"raw":612.2,"best":612.2,"ok":true,"reward":0.8721},{"n":2,"raw":753.3,"best":753.3,"ok":true,"reward":0.9579},{"n":3,"raw":774.6,"best":774.6,"ok":true,"reward":0.9708},{"n":4,"raw":829.2,"best":829.2,"ok":true,"reward":1.004},{"n":5,"raw":733.7,"best":829.2,"ok":true,"reward":0.9459},{"n":6,"raw":821.0,"best":829.2,"ok":true,"reward":0.999},{"n":7,"raw":976.5,"best":976.5,"ok":true,"reward":1.094},{"n":8,"raw":973.5,"best":976.5,"ok":true,"reward":1.092},{"n":9,"raw":960.3,"best":976.5,"ok":true,"reward":1.084}],"final_speedup":963.391397,"vs_oracle":1.171126},"deepseek":{"oracle":822.619615,"traj":[{"n":1,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":2,"raw":92.32,"best":92.32,"ok":true,"reward":0.5561},{"n":3,"raw":72.64,"best":92.32,"ok":true,"reward":0.5442},{"n":4,"raw":97.91,"best":97.91,"ok":true,"reward":0.5595},{"n":5,"raw":76.05,"best":97.91,"ok":true,"reward":0.5462},{"n":6,"raw":96.86,"best":97.91,"ok":true,"reward":0.5589},{"n":7,"raw":94.33,"best":97.91,"ok":true,"reward":0.5573},{"n":8,"raw":0.0,"best":97.91,"ok":true,"reward":0.0},{"n":9,"raw":0.0,"best":97.91,"ok":true,"reward":0.0},{"n":10,"raw":97.49,"best":97.91,"ok":true,"reward":0.5593},{"n":11,"raw":0.0,"best":97.91,"ok":true,"reward":0.0},{"n":12,"raw":97.02,"best":97.91,"ok":true,"reward":0.559}],"final_speedup":96.418413,"vs_oracle":0.117209}}},{"task_id":"wro-llamacpp-simd-q5k","short_name":"ggml Q5_K quantized dot-product AVX2 SIMD kernel","infra_name":"q5k_q8k_simd_dot","summary":"This is the CPU hot-path kernel ggml_vec_dot_q5_K_q8_K in the ggml tensor library used by llama.cpp: the dot product of a 5-bit K-quantized weight row with its paired activation row. The x86 implementation is slow; you must speed up this dot product with SIMD while reproducing results bit-for-bit. The input is paired Q5_K and Q8_K blocks, and the output is their dot product; specializing to the test inputs is not allowed. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 11.53): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"D","medium_topic":"D6","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":11.53,"traj":[{"n":1,"raw":10.37,"best":10.37,"ok":true,"reward":10.37},{"n":2,"raw":12.43,"best":12.43,"ok":true,"reward":12.43},{"n":3,"raw":11.97,"best":12.43,"ok":true,"reward":11.97},{"n":4,"raw":11.64,"best":12.43,"ok":true,"reward":11.64},{"n":5,"raw":11.88,"best":12.43,"ok":true,"reward":11.88}],"final_speedup":11.771568,"vs_oracle":1.020951},"kimi-k3":{"oracle":11.53,"traj":[{"n":1,"raw":12.11,"best":12.11,"ok":true,"reward":12.11},{"n":2,"raw":11.94,"best":12.11,"ok":true,"reward":11.94},{"n":3,"raw":11.92,"best":12.11,"ok":true,"reward":11.92},{"n":4,"raw":12.02,"best":12.11,"ok":true,"reward":12.02},{"n":5,"raw":11.23,"best":12.11,"ok":true,"reward":11.23},{"n":6,"raw":11.7,"best":12.11,"ok":true,"reward":11.7},{"n":7,"raw":11.53,"best":12.11,"ok":true,"reward":11.53}],"final_speedup":11.861091,"vs_oracle":1.028716},"gpt-5.6":{"oracle":11.53,"traj":[{"n":8,"raw":11.99,"best":11.99,"ok":true,"reward":11.99},{"n":16,"raw":11.43,"best":11.99,"ok":true,"reward":11.43}],"final_speedup":11.399695,"vs_oracle":0.988699},"glm-5.2":{"oracle":11.53,"traj":[{"n":1,"raw":11.72,"best":11.72,"ok":true,"reward":11.72},{"n":2,"raw":11.58,"best":11.72,"ok":true,"reward":11.58},{"n":3,"raw":11.72,"best":11.72,"ok":true,"reward":11.72}],"final_speedup":11.792347,"vs_oracle":1.022753},"sonnet-5":{"oracle":11.53,"traj":[{"n":1,"raw":5.623,"best":5.623,"ok":true,"reward":5.623},{"n":2,"raw":9.616,"best":9.616,"ok":true,"reward":9.616},{"n":3,"raw":10.24,"best":10.24,"ok":true,"reward":10.24},{"n":4,"raw":7.988,"best":10.24,"ok":true,"reward":7.988},{"n":5,"raw":10.65,"best":10.65,"ok":true,"reward":10.65},{"n":6,"raw":0.0,"best":10.65,"ok":false,"reward":0.0},{"n":7,"raw":8.622,"best":10.65,"ok":true,"reward":8.622},{"n":8,"raw":10.85,"best":10.85,"ok":true,"reward":10.85},{"n":9,"raw":11.93,"best":11.93,"ok":true,"reward":11.93},{"n":10,"raw":11.89,"best":11.93,"ok":true,"reward":11.89},{"n":11,"raw":12.04,"best":12.04,"ok":true,"reward":12.04},{"n":12,"raw":11.07,"best":12.04,"ok":true,"reward":11.07}],"final_speedup":11.864121,"vs_oracle":1.028978},"qwen3.7":{"oracle":11.53,"traj":[{"n":1,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":2,"raw":0.0,"best":null,"ok":false,"reward":0.0},{"n":3,"raw":6.76,"best":6.76,"ok":true,"reward":6.76},{"n":4,"raw":0.0,"best":6.76,"ok":false,"reward":0.0},{"n":5,"raw":6.937,"best":6.937,"ok":true,"reward":6.937},{"n":6,"raw":0.0,"best":6.937,"ok":false,"reward":0.0},{"n":7,"raw":6.62,"best":6.937,"ok":true,"reward":6.62},{"n":8,"raw":6.865,"best":6.937,"ok":true,"reward":6.865},{"n":9,"raw":7.977,"best":7.977,"ok":true,"reward":7.977},{"n":10,"raw":8.006,"best":8.006,"ok":true,"reward":8.006},{"n":11,"raw":7.817,"best":8.006,"ok":true,"reward":7.817},{"n":12,"raw":7.933,"best":8.006,"ok":true,"reward":7.933},{"n":13,"raw":7.064,"best":8.006,"ok":true,"reward":7.064},{"n":14,"raw":7.944,"best":8.006,"ok":true,"reward":7.944},{"n":15,"raw":7.348,"best":8.006,"ok":true,"reward":7.348}],"final_speedup":0.0,"vs_oracle":0.0},"qwen3.8":{"oracle":11.53,"traj":[{"n":1,"raw":12.3,"best":12.3,"ok":true,"reward":12.3},{"n":2,"raw":12.74,"best":12.74,"ok":true,"reward":12.74},{"n":3,"raw":12.17,"best":12.74,"ok":true,"reward":12.17},{"n":4,"raw":12.87,"best":12.87,"ok":true,"reward":12.87},{"n":5,"raw":11.97,"best":12.87,"ok":true,"reward":11.97},{"n":6,"raw":10.55,"best":12.87,"ok":true,"reward":10.55},{"n":7,"raw":11.37,"best":12.87,"ok":true,"reward":11.37},{"n":8,"raw":12.3,"best":12.87,"ok":true,"reward":12.3},{"n":9,"raw":12.25,"best":12.87,"ok":true,"reward":12.25},{"n":10,"raw":12.17,"best":12.87,"ok":true,"reward":12.17},{"n":11,"raw":11.91,"best":12.87,"ok":true,"reward":11.91},{"n":12,"raw":12.3,"best":12.87,"ok":true,"reward":12.3},{"n":13,"raw":12.18,"best":12.87,"ok":true,"reward":12.18},{"n":14,"raw":10.74,"best":12.87,"ok":true,"reward":10.74}],"final_speedup":12.027748,"vs_oracle":null},"deepseek":{"oracle":11.53,"traj":[{"n":1,"raw":11.11,"best":11.11,"ok":true,"reward":11.11},{"n":2,"raw":11.09,"best":11.11,"ok":true,"reward":11.09},{"n":3,"raw":0.0,"best":11.11,"ok":false,"reward":0.0},{"n":4,"raw":10.92,"best":11.11,"ok":true,"reward":10.92},{"n":5,"raw":10.9,"best":11.11,"ok":true,"reward":10.9},{"n":6,"raw":11.14,"best":11.14,"ok":true,"reward":11.14},{"n":7,"raw":11.0,"best":11.14,"ok":true,"reward":11.0},{"n":8,"raw":10.87,"best":11.14,"ok":true,"reward":10.87},{"n":9,"raw":11.09,"best":11.14,"ok":true,"reward":11.09},{"n":10,"raw":11.15,"best":11.15,"ok":true,"reward":11.15}],"final_speedup":11.071648,"vs_oracle":0.960247}}},{"task_id":"wro-fla-ttt-chunkscan-sol","short_name":"TTT (Test-Time Training) linear-attn chunked-scan kernel","infra_name":"ttt_linear_chunk_scan","summary":"This is chunk_ttt_linear, the test-time-training (TTT) linear layer in flash-linear-attention. The existing implementation is correct but slow; you must make it fast: the sequence is processed in mini-batches, and each batch performs one inner-loop gradient update to the fast-weight state under a layer-norm reconstruction objective, then reads out. The input is per-mini-batch q/k/v and parameters such as the learning rate, and the output is the result after the updated state and the output layer norm. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 63.1736): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"D","medium_topic":"D5","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":63.173556,"traj":[{"n":1,"raw":2.189,"best":2.189,"ok":true,"reward":0.5173},{"n":2,"raw":32.42,"best":32.42,"ok":true,"reward":0.7566},{"n":3,"raw":57.08,"best":57.08,"ok":true,"reward":0.9518},{"n":4,"raw":0.0,"best":57.08,"ok":false,"reward":0.0},{"n":5,"raw":59.84,"best":59.84,"ok":true,"reward":0.9737},{"n":6,"raw":60.94,"best":60.94,"ok":true,"reward":0.9823},{"n":7,"raw":62.22,"best":62.22,"ok":true,"reward":0.9925}],"final_speedup":64.106299,"vs_oracle":1.014765},"kimi-k3":{"oracle":63.173556,"traj":[{"n":1,"raw":21.84,"best":21.84,"ok":true,"reward":0.6729}],"final_speedup":65.910661,"vs_oracle":1.043327},"gpt-5.6":{"oracle":63.173556,"traj":[{"n":10,"raw":3.127,"best":3.127,"ok":true,"reward":0.5247},{"n":13,"raw":0.0,"best":3.127,"ok":true,"reward":0.0},{"n":14,"raw":0.0,"best":3.127,"ok":true,"reward":0.0}],"final_speedup":2.832253,"vs_oracle":0.044833},"glm-5.2":{"oracle":63.173556,"traj":[{"n":1,"raw":8.174,"best":8.174,"ok":true,"reward":0.5647},{"n":2,"raw":1.285,"best":8.174,"ok":true,"reward":0.5102},{"n":3,"raw":17.73,"best":17.73,"ok":true,"reward":0.6404},{"n":4,"raw":29.44,"best":29.44,"ok":true,"reward":0.733}],"final_speedup":29.373665,"vs_oracle":0.464968},"sonnet-5":{"oracle":63.173556,"traj":[{"n":1,"raw":1.408,"best":1.408,"ok":true,"reward":0.5111},{"n":2,"raw":2.406,"best":2.406,"ok":true,"reward":0.519},{"n":3,"raw":3.003,"best":3.003,"ok":true,"reward":0.5238},{"n":4,"raw":5.271,"best":5.271,"ok":true,"reward":0.5417},{"n":5,"raw":9.587,"best":9.587,"ok":true,"reward":0.5759},{"n":6,"raw":9.543,"best":9.587,"ok":true,"reward":0.5755},{"n":7,"raw":10.44,"best":10.44,"ok":true,"reward":0.5827}],"final_speedup":10.452368,"vs_oracle":0.165455},"qwen3.7":{"oracle":63.173556,"traj":[{"n":1,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":2,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":3,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":4,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":5,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":6,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":7,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":8,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":9,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":10,"raw":0.0,"best":0.0,"ok":true,"reward":0.08335},{"n":11,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":12,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":13,"raw":0.0,"best":0.0,"ok":true,"reward":0.0},{"n":14,"raw":0.0,"best":0.0,"ok":true,"reward":0.0}],"final_speedup":0.0,"vs_oracle":0.0},"qwen3.8":{"oracle":63.173556,"traj":[{"n":1,"raw":42.46,"best":42.46,"ok":true,"reward":0.8361},{"n":2,"raw":58.47,"best":58.47,"ok":true,"reward":0.9627},{"n":3,"raw":59.99,"best":59.99,"ok":true,"reward":0.9748},{"n":4,"raw":59.9,"best":59.99,"ok":true,"reward":0.9741},{"n":5,"raw":65.03,"best":65.03,"ok":true,"reward":1.015},{"n":6,"raw":65.2,"best":65.2,"ok":true,"reward":1.016},{"n":7,"raw":66.66,"best":66.66,"ok":true,"reward":1.028},{"n":8,"raw":70.56,"best":70.56,"ok":true,"reward":1.058},{"n":9,"raw":70.96,"best":70.96,"ok":true,"reward":1.062},{"n":10,"raw":70.19,"best":70.96,"ok":true,"reward":1.056},{"n":11,"raw":70.81,"best":70.96,"ok":true,"reward":1.06}],"final_speedup":71.218074,"vs_oracle":1.12734},"deepseek":{"oracle":63.173556,"traj":[{"n":1,"raw":0.9735,"best":0.9735,"ok":true,"reward":0.5077},{"n":2,"raw":1.891,"best":1.891,"ok":true,"reward":0.515},{"n":3,"raw":7.185,"best":7.185,"ok":true,"reward":0.5569},{"n":4,"raw":8.836,"best":8.836,"ok":true,"reward":0.5699},{"n":5,"raw":8.818,"best":8.836,"ok":true,"reward":0.5698},{"n":6,"raw":8.189,"best":8.836,"ok":true,"reward":0.5648},{"n":7,"raw":8.746,"best":8.836,"ok":true,"reward":0.5692},{"n":8,"raw":0.0,"best":8.836,"ok":true,"reward":0.3333},{"n":9,"raw":8.036,"best":8.836,"ok":true,"reward":0.5636}],"final_speedup":8.549671,"vs_oracle":0.135336}}},{"task_id":"wro-blocksparse-gemm-sol","short_name":"Block-sparse GEMM Triton kernel (nnz blocks only)","infra_name":"kblock_sparse_fp16_matmul","summary":"This is the matrix-multiply subsystem for structured K-block-sparse weights times fp16 activations: the K rows are split into several contiguous row blocks, only some blocks are non-zero (input-feature block pruning), and the weights are stored in compressed form. The existing blocksp_matmul is correct but slow; you must make it fast, and densifying the weights is not allowed. The input is the activations a, the compressed weight blocks w_blocks, the block indices k_idx, and block_k, and the output is a@W. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 26.5604): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"C","medium_topic":"C2","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":26.560385,"traj":[{"n":1,"raw":27.83,"best":27.83,"ok":true,"reward":1.024},{"n":2,"raw":45.04,"best":45.04,"ok":true,"reward":1.348},{"n":3,"raw":0.0,"best":45.04,"ok":false,"reward":0.0},{"n":4,"raw":40.99,"best":45.04,"ok":true,"reward":1.272},{"n":5,"raw":46.21,"best":46.21,"ok":true,"reward":1.37},{"n":6,"raw":46.2,"best":46.21,"ok":true,"reward":1.37},{"n":7,"raw":46.0,"best":46.21,"ok":true,"reward":1.366},{"n":8,"raw":46.89,"best":46.89,"ok":true,"reward":1.383},{"n":9,"raw":47.89,"best":47.89,"ok":true,"reward":1.402},{"n":10,"raw":46.73,"best":47.89,"ok":true,"reward":1.38},{"n":12,"raw":46.5,"best":47.89,"ok":true,"reward":1.375},{"n":13,"raw":46.54,"best":47.89,"ok":true,"reward":1.376},{"n":14,"raw":48.98,"best":48.98,"ok":true,"reward":1.422}],"final_speedup":48.215191,"vs_oracle":1.815305},"kimi-k3":{"oracle":26.560385,"traj":[{"n":1,"raw":11.43,"best":11.43,"ok":true,"reward":0.7151},{"n":2,"raw":21.2,"best":21.2,"ok":true,"reward":0.8992},{"n":3,"raw":39.46,"best":39.46,"ok":true,"reward":1.243},{"n":4,"raw":40.76,"best":40.76,"ok":true,"reward":1.267},{"n":5,"raw":44.17,"best":44.17,"ok":true,"reward":1.331},{"n":6,"raw":44.06,"best":44.17,"ok":true,"reward":1.329},{"n":7,"raw":44.25,"best":44.25,"ok":true,"reward":1.333},{"n":8,"raw":14.5,"best":44.25,"ok":true,"reward":0.773},{"n":9,"raw":43.93,"best":44.25,"ok":true,"reward":1.327}],"final_speedup":44.713492,"vs_oracle":1.683465},"gpt-5.6":{"oracle":26.560385,"traj":[{"n":15,"raw":23.11,"best":23.11,"ok":true,"reward":0.935}],"final_speedup":21.667742,"vs_oracle":0.815792},"glm-5.2":{"oracle":26.560385,"traj":[{"n":1,"raw":2.495,"best":2.495,"ok":true,"reward":0.547},{"n":2,"raw":2.836,"best":2.836,"ok":true,"reward":0.5534},{"n":3,"raw":4.087,"best":4.087,"ok":true,"reward":0.5769},{"n":4,"raw":5.784,"best":5.784,"ok":true,"reward":0.6089},{"n":5,"raw":5.814,"best":5.814,"ok":true,"reward":0.6094},{"n":6,"raw":5.783,"best":5.814,"ok":true,"reward":0.6089},{"n":7,"raw":5.253,"best":5.814,"ok":true,"reward":0.5989}],"final_speedup":5.894044,"vs_oracle":0.221911},"sonnet-5":{"oracle":26.560385,"traj":[{"n":1,"raw":12.74,"best":12.74,"ok":true,"reward":0.7398},{"n":2,"raw":27.04,"best":27.04,"ok":true,"reward":1.009},{"n":3,"raw":28.69,"best":28.69,"ok":true,"reward":1.04},{"n":4,"raw":28.01,"best":28.69,"ok":true,"reward":1.027},{"n":5,"raw":31.9,"best":31.9,"ok":true,"reward":1.101},{"n":6,"raw":4.778,"best":31.9,"ok":true,"reward":0.5899},{"n":7,"raw":31.13,"best":31.9,"ok":true,"reward":1.086}],"final_speedup":29.984386,"vs_oracle":1.128914},"qwen3.7":{"oracle":26.560385,"traj":[{"n":1,"raw":1.241,"best":1.241,"ok":true,"reward":0.5234},{"n":2,"raw":0.0,"best":1.241,"ok":false,"reward":0.0},{"n":3,"raw":1.283,"best":1.283,"ok":true,"reward":0.5242},{"n":4,"raw":0.0,"best":1.283,"ok":false,"reward":0.0},{"n":5,"raw":2.106,"best":2.106,"ok":true,"reward":0.5396},{"n":6,"raw":2.108,"best":2.108,"ok":true,"reward":0.5397},{"n":7,"raw":2.894,"best":2.894,"ok":true,"reward":0.5545},{"n":8,"raw":2.767,"best":2.894,"ok":true,"reward":0.5521},{"n":9,"raw":3.595,"best":3.595,"ok":true,"reward":0.5677},{"n":10,"raw":2.862,"best":3.595,"ok":true,"reward":0.5539},{"n":11,"raw":2.733,"best":3.595,"ok":true,"reward":0.5514},{"n":12,"raw":2.793,"best":3.595,"ok":true,"reward":0.5526},{"n":13,"raw":1.873,"best":3.595,"ok":true,"reward":0.5353},{"n":14,"raw":2.789,"best":3.595,"ok":true,"reward":0.5525},{"n":15,"raw":2.745,"best":3.595,"ok":true,"reward":0.5517},{"n":16,"raw":2.731,"best":3.595,"ok":true,"reward":0.5514}],"final_speedup":2.743199,"vs_oracle":0.103282},"qwen3.8":{"oracle":26.560385,"traj":[{"n":1,"raw":20.67,"best":20.67,"ok":true,"reward":0.8891},{"n":2,"raw":34.94,"best":34.94,"ok":true,"reward":1.158},{"n":3,"raw":44.6,"best":44.6,"ok":true,"reward":1.34}],"final_speedup":44.100625,"vs_oracle":1.660391},"deepseek":{"oracle":26.560385,"traj":[{"n":1,"raw":0.1768,"best":0.1768,"ok":true,"reward":0.5033},{"n":2,"raw":0.1642,"best":0.1768,"ok":true,"reward":0.5031},{"n":3,"raw":30.32,"best":30.32,"ok":true,"reward":1.071},{"n":4,"raw":20.77,"best":30.32,"ok":true,"reward":0.8909},{"n":5,"raw":21.62,"best":30.32,"ok":true,"reward":0.907},{"n":6,"raw":30.43,"best":30.43,"ok":true,"reward":1.073},{"n":7,"raw":27.49,"best":30.43,"ok":true,"reward":1.018},{"n":8,"raw":29.17,"best":30.43,"ok":true,"reward":1.049},{"n":9,"raw":29.82,"best":30.43,"ok":true,"reward":1.061},{"n":10,"raw":24.08,"best":30.43,"ok":true,"reward":0.9534},{"n":11,"raw":0.0,"best":30.43,"ok":true,"reward":0.0},{"n":12,"raw":29.57,"best":30.43,"ok":true,"reward":1.057},{"n":13,"raw":29.22,"best":30.43,"ok":true,"reward":1.05},{"n":14,"raw":16.11,"best":30.43,"ok":true,"reward":0.8033},{"n":15,"raw":30.57,"best":30.57,"ok":true,"reward":1.076}],"final_speedup":29.104345,"vs_oracle":1.09578}}},{"task_id":"wro-flashattn-cute-blocksparse","short_name":"FlashAttention CuTe block-sparse forward","infra_name":"blocksparse_attn_forward","summary":"This is the forward of block-sparse multi-head attention in the flash-attention CuTe version: a query may only attend to keys within the allowed (query-block, key-block) pairs. The existing implementation is correct but slow; you must make the forward fast, and neither the signature nor the block-sparse mask semantics may be changed. The input is q/k/v and the description of the allowed blocks, and the output is the block-sparse attention output. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 189.21): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"D","medium_topic":"D1","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":189.21,"traj":[{"n":1,"raw":160.5,"best":160.5,"ok":true,"reward":0.9242},{"n":2,"raw":176.0,"best":176.0,"ok":true,"reward":0.965},{"n":3,"raw":174.1,"best":176.0,"ok":true,"reward":0.9601},{"n":4,"raw":168.7,"best":176.0,"ok":true,"reward":0.9458},{"n":5,"raw":175.9,"best":176.0,"ok":true,"reward":0.9647},{"n":6,"raw":193.3,"best":193.3,"ok":true,"reward":1.011},{"n":7,"raw":191.7,"best":193.3,"ok":true,"reward":1.006}],"final_speedup":192.810727,"vs_oracle":1.01903},"kimi-k3":{"oracle":189.21,"traj":[{"n":1,"raw":188.9,"best":188.9,"ok":true,"reward":0.9992},{"n":2,"raw":189.1,"best":189.1,"ok":true,"reward":0.9997},{"n":3,"raw":190.7,"best":190.7,"ok":true,"reward":1.004},{"n":4,"raw":197.4,"best":197.4,"ok":true,"reward":1.022},{"n":5,"raw":190.7,"best":197.4,"ok":true,"reward":1.004},{"n":6,"raw":190.4,"best":197.4,"ok":true,"reward":1.003},{"n":7,"raw":199.0,"best":199.0,"ok":true,"reward":1.026},{"n":8,"raw":197.6,"best":199.0,"ok":true,"reward":1.022}],"final_speedup":195.383668,"vs_oracle":1.032629},"gpt-5.6":{"oracle":189.21,"traj":[{"n":9,"raw":183.6,"best":183.6,"ok":true,"reward":0.9853}],"final_speedup":183.220071,"vs_oracle":0.968342},"glm-5.2":{"oracle":189.21,"traj":[{"n":1,"raw":12.97,"best":12.97,"ok":true,"reward":0.5343},{"n":2,"raw":17.15,"best":17.15,"ok":true,"reward":0.5453},{"n":3,"raw":17.85,"best":17.85,"ok":true,"reward":0.5472},{"n":4,"raw":20.52,"best":20.52,"ok":true,"reward":0.5542},{"n":5,"raw":22.05,"best":22.05,"ok":true,"reward":0.5583}],"final_speedup":22.021495,"vs_oracle":0.116387},"sonnet-5":{"oracle":189.21,"traj":[{"n":1,"raw":109.7,"best":109.7,"ok":true,"reward":0.7898},{"n":2,"raw":110.9,"best":110.9,"ok":true,"reward":0.7931},{"n":3,"raw":123.5,"best":123.5,"ok":true,"reward":0.8263},{"n":4,"raw":127.1,"best":127.1,"ok":true,"reward":0.836},{"n":5,"raw":126.4,"best":127.1,"ok":true,"reward":0.8341}],"final_speedup":126.192912,"vs_oracle":0.666946},"qwen3.7":{"oracle":189.21,"traj":[{"n":1,"raw":190.2,"best":190.2,"ok":true,"reward":1.003},{"n":2,"raw":188.2,"best":190.2,"ok":true,"reward":0.9972},{"n":3,"raw":190.3,"best":190.3,"ok":true,"reward":1.003}],"final_speedup":189.500213,"vs_oracle":1.001534},"qwen3.8":{"oracle":189.21,"traj":[{"n":1,"raw":65.9,"best":65.9,"ok":true,"reward":0.6741},{"n":2,"raw":65.48,"best":65.9,"ok":true,"reward":0.673},{"n":3,"raw":172.7,"best":172.7,"ok":true,"reward":0.9564},{"n":4,"raw":182.3,"best":182.3,"ok":true,"reward":0.9817},{"n":5,"raw":185.6,"best":185.6,"ok":true,"reward":0.9906},{"n":6,"raw":183.6,"best":185.6,"ok":true,"reward":0.9851}],"final_speedup":182.355646,"vs_oracle":0.963774},"deepseek":{"oracle":189.21,"traj":[{"n":1,"raw":1.807,"best":1.807,"ok":true,"reward":0.5048},{"n":2,"raw":0.0,"best":1.807,"ok":true,"reward":0.0},{"n":3,"raw":0.0,"best":1.807,"ok":true,"reward":0.4167},{"n":4,"raw":2.274,"best":2.274,"ok":true,"reward":0.506},{"n":5,"raw":2.285,"best":2.285,"ok":true,"reward":0.506},{"n":6,"raw":0.0,"best":2.285,"ok":true,"reward":0.4722},{"n":7,"raw":2.082,"best":2.285,"ok":true,"reward":0.5055}],"final_speedup":2.212235,"vs_oracle":0.011692}}},{"task_id":"wro-flashattn-cute-scoremod-tanh","short_name":"wro-flashattn-cute-scoremod-tanh","infra_name":"attn_score_mod_tanh","summary":"This is the attention forward with a user-defined score modification in the flash-attention CuTe version: before softmax, a callable transform is applied to each scaled score. The existing implementation is correct but slow; you must make this gated-tanh-form score-mod path fast (it may include a causal mask and grouped-query attention). The input is q/k/v and the score transform, and the output is the attention output. Metric: continuous performance-class score — first pass a hard correctness gate, then score on a logarithmic curve relative to the oracle anchor (ref_speedup = 31.37): tying the anchor scores 0, you must exceed it to score anything, and reaching the anchor squared caps at 1.0. The per-round dev progress signal reported to the solver is `dev_speedup` (the measured baseline_ms / candidate_ms wall-clock ratio), and the scored `speedup` fed into the curve is this same ratio measured by the held-out verifier.","bench":"lh","big_topic":"D","medium_topic":"D1","task_class":"performance","is_performance":true,"metric":"speedup","lower_better":false,"models":{"opus-5":{"oracle":31.37,"traj":[{"n":1,"raw":19.68,"best":19.68,"ok":true,"reward":0.8136},{"n":2,"raw":25.08,"best":25.08,"ok":true,"reward":0.8997},{"n":3,"raw":25.61,"best":25.61,"ok":true,"reward":0.9082}],"final_speedup":25.467085,"vs_oracle":0.811829},"kimi-k3":{"oracle":31.37,"traj":[{"n":1,"raw":31.86,"best":31.86,"ok":true,"reward":1.008},{"n":2,"raw":31.78,"best":31.86,"ok":true,"reward":1.007},{"n":3,"raw":31.79,"best":31.86,"ok":true,"reward":1.007},{"n":4,"raw":31.85,"best":31.86,"ok":true,"reward":1.008}],"final_speedup":31.790388,"vs_oracle":1.013401},"gpt-5.6":{"oracle":31.37,"traj":[{"n":1,"raw":30.84,"best":30.84,"ok":true,"reward":0.9915}],"final_speedup":30.845128,"vs_oracle":0.983268},"glm-5.2":{"oracle":31.37,"traj":[{"n":1,"raw":31.5,"best":31.5,"ok":true,"reward":1.002},{"n":2,"raw":31.52,"best":31.52,"ok":true,"reward":1.002}],"final_speedup":31.518861,"vs_oracle":1.004745},"sonnet-5":{"oracle":31.37,"traj":[{"n":1,"raw":31.49,"best":31.49,"ok":true,"reward":1.002},{"n":2,"raw":0.0,"best":31.49,"ok":true,"reward":0.4722},{"n":3,"raw":31.5,"best":31.5,"ok":true,"reward":1.002}],"final_speedup":31.461093,"vs_oracle":1.002904},"qwen3.7":{"oracle":31.37,"traj":[{"n":1,"raw":31.38,"best":31.38,"ok":true,"reward":1.0},{"n":2,"raw":31.31,"best":31.38,"ok":true,"reward":0.999},{"n":3,"raw":31.39,"best":31.39,"ok":true,"reward":1.0}],"final_speedup":31.396425,"vs_oracle":1.000842},"qwen3.8":{"oracle":31.37,"traj":[{"n":1,"raw":17.91,"best":17.91,"ok":true,"reward":0.7855},{"n":2,"raw":19.2,"best":19.2,"ok":true,"reward":0.806},{"n":3,"raw":22.96,"best":22.96,"ok":true,"reward":0.8659},{"n":4,"raw":22.89,"best":22.96,"ok":true,"reward":0.8649},{"n":5,"raw":22.99,"best":22.99,"ok":true,"reward":0.8664},{"n":6,"raw":24.91,"best":24.91,"ok":true,"reward":0.8971},{"n":7,"raw":24.98,"best":24.98,"ok":true,"reward":0.8982},{"n":8,"raw":25.01,"best":25.01,"ok":true,"reward":0.8986}],"final_speedup":25.073618,"vs_oracle":0.799287},"deepseek":{"oracle":31.37,"traj":[{"n":1,"raw":31.45,"best":31.45,"ok":true,"reward":1.001}],"final_speedup":31.480792,"vs_oracle":1.003532}}},{"task_id":"wli-fla-kkt-solvetril-build","short_name":"wli-fla-kkt-solvetril-build","infra_name":"wy_ut_transform_solve_tril","summary":"This is the intra-chunk WY/UT transform in fla-org/flash-linear-attention (a flagship open-source LLM-infra kernel library), the intra-chunk mixing primitive for delta-rule and gated-delta linear attention. The implementation across two coupled files is for you to complete: compute the A matrix and solve T=(I+A)^{-1} (lower-triangular inverse); this task is judged only on correctness. The input is per-chunk k, beta, and the gate g, and the output is A and T, which must chain correctly into the downstream delta_rule / gated_delta_product. Metric: binary implementation-class score — 1.0 only if all hidden cases and hard gates pass; any single case failure or a triggered forbidden-edit/cheat condition is 0.0, with no partial credit.","bench":"lh","big_topic":"I","medium_topic":"I1","task_class":"implementation","is_performance":false,"metric":"speedup","lower_better":false,"models":{"qwen3.8":{"oracle":null,"traj":[{"n":1,"raw":0.0,"best":0.0,"ok":true,"reward":1.0}],"final_speedup":null,"vs_oracle":null}}}]}}