·
AI & ML interests
Contact: arxivgpt@gmail.com
Recent Activity
reacted to theirpost with 😎 about 21 hours ago We wrote up our run in The Fast Gemma Challenge — as vidraft-darwin — and wanted to share the recipe. 🙏
https://huggingface.co/spaces/gemma-challenge/gemma-dashboard
Verified result: 510.58 TPS at PPL 2.3930 on a single A10G (fw188-ctk49-n64-patchbridge, re-run & VERIFIED). Honest note: on raw TPS there are faster runs (535+), but those went over the PPL bar and didn't verify — what we're proud of is the fastest result that keeps quality.
The recipe is already open, so we explained each piece: sliding-window W188, CTK49 kernel tuning, noprecache (honest, verifiable measurement), and an N64 synthetic warmup bridge that shrinks the public↔private gap (~15 TPS), plus INT4 + MTP K=7 + CUDA-graph capture. One rule: only stack quality-neutral speedups.
Huge thanks to @firfir-cast, @gemma-slayer, @chiku-inu, @kenyan-duma, @dixie-flatline and everyone who shared their experiments. Full write-up
👇
https://huggingface.co/blog/FINAL-Bench/fast-gemma
reacted to theirpost with 🧠 about 21 hours ago We wrote up our run in The Fast Gemma Challenge — as vidraft-darwin — and wanted to share the recipe. 🙏
https://huggingface.co/spaces/gemma-challenge/gemma-dashboard
Verified result: 510.58 TPS at PPL 2.3930 on a single A10G (fw188-ctk49-n64-patchbridge, re-run & VERIFIED). Honest note: on raw TPS there are faster runs (535+), but those went over the PPL bar and didn't verify — what we're proud of is the fastest result that keeps quality.
The recipe is already open, so we explained each piece: sliding-window W188, CTK49 kernel tuning, noprecache (honest, verifiable measurement), and an N64 synthetic warmup bridge that shrinks the public↔private gap (~15 TPS), plus INT4 + MTP K=7 + CUDA-graph capture. One rule: only stack quality-neutral speedups.
Huge thanks to @firfir-cast, @gemma-slayer, @chiku-inu, @kenyan-duma, @dixie-flatline and everyone who shared their experiments. Full write-up
👇
https://huggingface.co/blog/FINAL-Bench/fast-gemma
reacted to theirpost with 👍 about 21 hours ago We wrote up our run in The Fast Gemma Challenge — as vidraft-darwin — and wanted to share the recipe. 🙏
https://huggingface.co/spaces/gemma-challenge/gemma-dashboard
Verified result: 510.58 TPS at PPL 2.3930 on a single A10G (fw188-ctk49-n64-patchbridge, re-run & VERIFIED). Honest note: on raw TPS there are faster runs (535+), but those went over the PPL bar and didn't verify — what we're proud of is the fastest result that keeps quality.
The recipe is already open, so we explained each piece: sliding-window W188, CTK49 kernel tuning, noprecache (honest, verifiable measurement), and an N64 synthetic warmup bridge that shrinks the public↔private gap (~15 TPS), plus INT4 + MTP K=7 + CUDA-graph capture. One rule: only stack quality-neutral speedups.
Huge thanks to @firfir-cast, @gemma-slayer, @chiku-inu, @kenyan-duma, @dixie-flatline and everyone who shared their experiments. Full write-up
👇
https://huggingface.co/blog/FINAL-Bench/fast-gemma
View all activity Organizations
view article The Fast Gemma Challenge: our verified-SOTA recipe, in full
FINAL-Bench
• • 17
view article POCKET: a 35-billion-parameter model that runs on your iPhone — and on your PC with no GPU
FINAL-Bench
• • 12
view article Aether-7B-5Attn: A 100% Open-Source Sovereign Foundation Model — and a Controlled Experiment in Heterogeneous Attention
FINAL-Bench
• • 21
view article VKUE: No GPU? Runs Anyway — a 34.7B Reasoner on a Laptop and on Bare CPU
FINAL-Bench
• • 17
published an article about 1 month ago view article Quantum Cryptanalysis on Real Hardware: Pushing Symmetric-Structure Key Recovery Beyond the Published Frontier
FINAL-Bench
• • 15
published an article about 1 month ago published an article about 1 month ago view article Chitos: From Detection to Proof — An Autonomous Security AI That Actually Exploits
FINAL-Bench
• • 19
published an article about 2 months ago view article FINAL-Bench Quantum: An Open, Neutral Benchmark for Quantum-Computing Methods
FINAL-Bench
• • 17
view article Training-Free Reasoning at 88.89% on GPQA Diamond: How Darwin Family Hit Frontier Scores Without a Single Gradient Step
FINAL-Bench
• • 18
view article Darwin-TTS: We Gave a TTS Model 3% of an LLM's Brain — It Started Showing Emotion
FINAL-Bench
• • 13
view article "Darwin-27B-Opus: Surpassing the Foundation Model Without Training"
FINAL-Bench
• • 16
view article Darwin V6: Diagnostic-Guided Evolutionary Model Merging
view article "The Child That Surpassed Both Parents Through MRI-Guided Evolutionary Merge"
FINAL-Bench
• • 15
view article Introducing WM Bench: A Benchmark for Cognitive Intelligence in World Models
FINAL-Bench
• • 13
view article 🏟️ Smol AI WorldCup: A 5-Axis Benchmark That Reveals What Small Language Models Can Really Do
FINAL-Bench
• • 38
view article MARL: Runtime Middleware That Reduces LLM Hallucination Without Fine-Tuning
view article Structural Problems in AI Benchmarking and the Case for a Unified Evaluation Framework
view article Do Bubbles Form When Tens of Thousands of AIs Simulate Capitalism?
FINAL-Bench
• • 17
view article FINAL Bench: The Real Bottleneck to AGI Is Self-Correction
FINAL-Bench
• • 20