Go
Home
Pricing
FAQ
Install
Home
Pricing
FAQ
Install
Login
via GitHub
huggingface/lighteval
Pull Requests
Commits
Open
Closed
fix(ci): harden GitHub Actions workflows (#1388)
#1389 opened 2026-09-17 08:55 by
hf-security-analysis[bot]
Bump the actions group with 8 updates
dependencies
github_actions
#1388 opened 2026-09-17 08:54 by
dependabot[bot]
Fix logging call that hides why torch.compile failed
#1387 opened 2026-09-16 05:26 by
David-Wu1119
perf: batch-tokenize contexts/continuations in _loglikelihood_tokens
#1386 opened 2026-09-11 02:02 by
drmohanty-tech
fix: apply strip_strings/normalize in avg_at_n (preprocess was skipped)
#1385 opened 2026-09-10 15:20 by
Abelo9996
feat: select few-shot examples by dataset ids
#1381 opened 2026-09-08 07:02 by
foxboy66
fix: isolate few-shot sampler RNG and pool from shared and global state (#1307, #1309)
#1379 opened 2026-09-07 04:52 by
Abelo9996
Pass max_tokens as an int on the litellm judge backend
#1378 opened 2026-09-04 04:43 by
gyanu2507
Add unit tests for VLLMModel core methods (#724)
#1377 opened 2026-09-02 19:14 by
sneha4175
Fix normalized multiple-choice probability underflow
#1376 opened 2026-09-02 01:58 by
YusefSyed
fix(metrics): add actionable BERTScore baseline guidance
#1373 opened 2026-08-31 03:16 by
LarryHu0217
fix(tasks): skip few-shot reuse warning for zero-shot tasks
#1372 opened 2026-08-31 03:15 by
LarryHu0217
fix(judge): honor concurrent_requests for the OpenAI/TGI judge backend
#1371 opened 2026-08-30 23:51 by
mohansree14
Key results and details output paths on the model revision
#1370 opened 2026-08-30 04:02 by
thisisreyy
Pin prediction cache to resolved model revisions
#1369 opened 2026-08-30 02:47 by
LarryHu0217
fix(cache): make multi-split document IDs unique
#1368 opened 2026-08-30 01:38 by
linhongyu510
feat: add OrcaRouter as a first-class provider in the LiteLLM backend
#1367 opened 2026-08-30 01:31 by
kuswardhanietidims-svg
fix: resolve SamplingMetric string normalize argument
#1366 opened 2026-08-29 22:37 by
Abelo9996
fix: honor task-level num_samples and warn when it has no effect
#1361 opened 2026-08-27 20:39 by
Abelo9996
ci: bump pr_style_bot to current style-bot-action reusable workflow
#1360 opened 2026-08-27 20:29 by
Abelo9996
Fix Dyck scoring to preserve bracket order
#1358 opened 2026-08-25 15:30 by
Excelius-Wang
Fix continuation tokens after distributed gathering
#1357 opened 2026-08-25 13:06 by
Excelius-Wang
fix(sglang): size batch truncation from the longest input in the batch, not the first
#1355 opened 2026-08-24 10:20 by
shoemoney
don't score an empty F1 gold reference as zero
#1354 opened 2026-08-24 03:13 by
gyanu2507
docs: fix broken code examples in evaluating-a-custom-model.mdx
#1352 opened 2026-08-22 19:38 by
mohansree14
Fix duplicate custom task module imports
#1351 opened 2026-08-21 12:42 by
yangweigbh
fix: keep every task's responses when reloading them from details
#1350 opened 2026-08-19 19:33 by
tonycoder-hub
Fix relational comparisons with different condition counts
#1349 opened 2026-08-19 17:24 by
Excelius-Wang
Parse hellaswag_arabic endings with ast.literal_eval
#1348 opened 2026-08-19 08:11 by
Nakul-Sinha
Add add_generation_prompt=True when tokenizing judge prompts
#1347 opened 2026-08-19 08:11 by
Nakul-Sinha
Older