llama.cpp
llama-quantize: Add MoE chunk queue for faster multi-threaded quant creation
#27770
Open

llama-quantize: Add MoE chunk queue for faster multi-threaded quant creation #27770

bartowski1182 wants to merge 1 commit into ggml-org:master from bartowski1182:moe-chunk-queue
bartowski1182
bartowski1182 Add MoE chunk queue
1a744c9b
bartowski1182 bartowski1182 marked this pull request as ready for review 2 days ago
bartowski1182 bartowski1182 requested a review from ggerganov ggerganov 2 days ago
ngxson ngxson assigned ngxson ngxson 1 day ago
bartowski1182 bartowski1182 marked this pull request as draft 18 hours ago

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
Labels
Milestone