chat-ui
07ddbb22 - Set explicit max_tokens for DeepSeek-V4-Flash-0731 (#2482)

Commit
57 days ago
Set explicit max_tokens for DeepSeek-V4-Flash-0731 (#2482) The entry added in #2479 sent no max_tokens, so requests fell back to provider default output caps. Baseten (fastest provider for this model) defaults to 4096 output tokens with thinking enabled by default, so coding and artifact requests were truncated mid-answer and surfaced as normal completions. Align with the older DeepSeek-V4-Flash entry (49152) in prod and dev.
Author
Parents
Loading