Go
Home
Pricing
FAQ
Install
Home
Pricing
FAQ
Install
Login
via GitHub
huggingface/lighteval
Pull Requests
Commits
Open
Closed
Fix Metrics.drop corpus aggregation to use mean instead of max
#1422 opened 2026-10-09 16:55 by
salignatmoandal
Fix broken code examples in custom model docs
#1421 opened 2026-10-09 10:30 by
singularity-14
Reject a missing details file when reloading an eval
#1420 opened 2026-10-08 18:05 by
b423016
Show scored sample counts in the results table
#1419 opened 2026-10-08 18:02 by
b423016
fix: show a single vLLM data-parallel progress bar on the driver
#1418 opened 2026-10-07 13:52 by
Steeve-Crypto
fix(metrics): ignore boundary whitespace in word-unit counts
#1416 opened 2026-10-04 17:10 by
March-7
Don't change global torch matmul precision when importing swiss_legal metrics
#1415 opened 2026-10-02 06:53 by
jayzuccarelli
fix(tasks): bypass deprecated dataset scripts with parquet data files (#1406)
#1413 opened 2026-10-01 08:12 by
irvallensar
fix(metrics): raise when pass@k is called with k > n
#1412 opened 2026-10-01 06:32 by
woodforestsen
fix: litellm judge backend silently scores API failures as 0 instead of surfacing them
#1411 opened 2026-09-30 05:30 by
johnny-glitch12
Add sample count column to markdown results table
#1409 opened 2026-09-29 19:20 by
manan45
Fix DROP exact match treating a repeated span as one copy
#1408 opened 2026-09-29 15:08 by
SashaMIT
Fix LiveCodeBench scoring every generation as -2 under spawn and forkserver
#1401 opened 2026-09-27 22:49 by
David-Wu1119
math tasks: pass math_normalizer to maj@n as normalize
#1399 opened 2026-09-25 18:21 by
no-hup
Fix model arg parsing for endpoint URLs
#1398 opened 2026-09-25 17:23 by
saptarshihalder
fix(metrics): let pass_at_k_letters grade boxed / chain-of-thought answers (fixes #1094)
#1392 opened 2026-09-22 07:55 by
Linxiushen
docs: explain how to run built-in multilingual tasks
#1391 opened 2026-09-19 11:39 by
buyan201430-code
Allow configuring the LiteLLM cache directory
#1390 opened 2026-09-19 06:31 by
lindicaphxag-tech
Bump the actions group with 8 updates
dependencies
github_actions
#1388 opened 2026-09-17 08:54 by
dependabot[bot]
Fix logging call that hides why torch.compile failed
#1387 opened 2026-09-16 05:26 by
David-Wu1119
perf: batch-tokenize contexts/continuations in _loglikelihood_tokens
#1386 opened 2026-09-11 02:02 by
drmohanty-tech
fix: apply strip_strings/normalize in avg_at_n (preprocess was skipped)
#1385 opened 2026-09-10 15:20 by
Abelo9996
feat: select few-shot examples by dataset ids
#1381 opened 2026-09-08 07:02 by
foxboy66
fix: isolate few-shot sampler RNG and pool from shared and global state (#1307, #1309)
#1379 opened 2026-09-07 04:52 by
Abelo9996
Pass max_tokens as an int on the litellm judge backend
#1378 opened 2026-09-04 04:43 by
gyanu2507
Add unit tests for VLLMModel core methods (#724)
#1377 opened 2026-09-02 19:14 by
sneha4175
Fix normalized multiple-choice probability underflow
#1376 opened 2026-09-02 01:58 by
YusefSyed
fix(metrics): add actionable BERTScore baseline guidance
#1373 opened 2026-08-31 03:16 by
LarryHu0217
fix(tasks): skip few-shot reuse warning for zero-shot tasks
#1372 opened 2026-08-31 03:15 by
LarryHu0217
fix(judge): honor concurrent_requests for the OpenAI/TGI judge backend
#1371 opened 2026-08-30 23:51 by
mohansree14
Older