LocalMaxxing
Primeros pasos
Modelos
Hardware
Evaluaciones comparativas
Más
+
Enviar
Primeros pasos
Clasificación
Modelos
Hardware
Evaluaciones comparativas
Marketplace
Alquileres
Pro
Docs de API
Idioma
English
简体中文
繁體中文
日本語
한국어
Español
Français
Deutsch
Italiano
Português (Brasil)
Русский
Polski
Nederlands
Türkçe
हिन्दी
Bahasa Indonesia
Tiếng Việt
ไทย
Back to benchmarks
Manual
AI draft
Build an benchmark
Create a benchmark card, attach inline or bucket-backed datasets, and submit it for approval.
Name *
Slug *
Category *
Runner *
Custom local
LM-Eval Harness
Scoring *
Exact match
F1
Pass@k
LLM judge
User rating
Description
Source URL
Task 1
Task key *
Display name *
Weight
Task type
Multiple choice
QA
Code
Judge
Dataset source
Inline samples
S3 bucket upload
URL
Hugging Face
Prompt template
{{input}}
Sample 1
Input *
Choices, one per line *
Gold answer *
Add sample
Add another task