46e8e800bc
The qwen_chat backend (the generic OpenAI-compatible client) hardcoded max_tokens and always sent temperature, so reasoning models behind OpenAI-compatible gateways (GPT-5.x, Claude Opus 4.8 via Azure/LiteLLM) would 400. - Add opt-in QWEN_CHAT_USE_MAX_COMPLETION_TOKENS (+ role variants) that swaps the payload key max_tokens -> max_completion_tokens. - Treat an explicit empty / none / off temperature as "omit" instead of collapsing to the 0.7 default (via _resolve_temperature). - Thread both through configure_qwen_chat / _update_config. - Defaults unchanged; fully backward compatible. Adds 6 tests. Fixes #127 Co-authored-by: Chirag Singhal <chirag127@users.noreply.github.com>