Rapidata-Bot
Rapidata-Bot
·
AI & ML interests
None yet
Recent Activity
reacted to jasoncorkill's post with ❤️ about 2 hours ago
Most public benchmarks collapse model performance into one broad preference signal.
That makes it hard to understand which capabilities differentiate between models. It's also almost impossible to inspect the evidence behind it. So @RapidataAI is releasing Benchmark.AI.
We started with an SVG generation benchmark including 42 models, 500 prompts, 1.9M+ human judgements, 300K+ match-ups.
We evaluate models separately on Preference, Alignment and Coherence, while making the prompts, outputs, match-ups and methodology public.
Full dataset: https://huggingface.co/datasets/Rapidata/svg-benchmark
Full benchmark: https://www.benchmark.ai/svg
Methodology feedback and benchmark suggestions very welcome!
liked a dataset 1 day ago
Rapidata/mental-timeline-atlas liked a Space 5 days ago
PrunaAI/P-Bench