Proto_AGI's picture

Proto_AGI

mayafree

AI & ML interests

None yet

Recent Activity

liked a Space about 5 hours ago
FINAL-Bench/ONGRID
reacted to SeaWolf-AI's post with πŸ”₯ about 15 hours ago
🧠 We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything. Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route. βš™οΈ How it works πŸ”Ή It makes its call in a single forward pass. πŸ”Ή Zero generated tokens, and no decoding loop. πŸ”Ή That keeps latency and cost far below what a generative model needs. 🎯 What it judges πŸ”Ή It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score). πŸ”Ή For each one it hands back a calibrated confidence, not just an answer. πŸ“Š How well calibrated (measured) πŸ”Ή KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens. πŸ”Ή 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors. πŸ”Ή By type: noul 0.847, choice 0.723, score 0.675. πŸ”Ή None of the benchmark's train split went into it. It is pure zero-shot. πŸš€ Where it fits πŸ”Ή Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation. πŸ† It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot). πŸ”— Links Model: https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC Leaderboard: https://huggingface.co/datasets/LocalLLaMA/typed-decisions Curious to hear what you make of the single-pass, no-generation approach. πŸ™Œ
View all activity

Organizations

mayafree_ai's profile picture Gemma Challenge's profile picture