What happened?
Thanks for the usage-chunk handling in create_stream (#3972) and the normalized stop reasons (#5027). GPT-6 Sol, Luna and Astra on Amazon Bedrock already stream through OpenAIChatCompletionClient with just base_url, api_key and model_info set. One result field still differs from create().
Describe the bug
create_stream reads finish_reason only from chunks where usage is None (_openai_client.py L948). The Bedrock runtime endpoint /openai/v1/chat/completions puts usage on the GPT-6 finish chunk, with or without include_usage. So the finish reason is never read and the streamed result says "unknown", including for length truncation and tool calls. #4875 (DeepSeek, "ValueError: No stop reason found" always raised when llm usage returned) looks like the same code path; #5027 turned that error into "unknown".
| live, main@027ecf0a |
create() |
create_stream() today |
with the fix below |
runtime us-east-1 us.openai.gpt-6-sol and global.openai.gpt-6-sol: text / truncated / tool call¹ |
stop / length / function_calls |
❌ unknown ×3 |
✅ stop / length / function_calls |
runtime us-east-1 us.openai.gpt-6-luna and global.openai.gpt-6-luna: same 3 |
same |
❌ unknown ×3 |
✅ same as create() |
runtime us.openai.gpt-6-astra (us-east-1, us-west-2) and global.openai.gpt-6-astra (us-east-1): text / truncated |
stop / length |
❌ unknown ×2 |
✅ stop / length |
runtime openai.gpt-oss-120b-1:0 (no usage on finish chunk) |
stop / length / function_calls |
✅ same as create() |
✅ unchanged |
Mantle openai.gpt-6-sol (us-east-1, 3 cases) and openai.gpt-6-astra (us-west-2, text / truncated), no usage on finish chunk |
stop / length (/ function_calls) |
✅ same as create() |
✅ unchanged |
| streamed usage, GPT-6 Sol text |
— |
(8, 5) |
✅ (8, 5) unchanged |
¹ with reasoning_effort="none", which Bedrock chat completions needs for Sol/Luna function tools.
To Reproduce
import asyncio
import os
from autogen_core.models import ModelInfo, UserMessage
from autogen_ext.models.openai import OpenAIChatCompletionClient
async def main() -> None:
for model in ["us.openai.gpt-6-sol", "us.openai.gpt-6-luna", "us.openai.gpt-6-astra"]:
client = OpenAIChatCompletionClient(
model=model,
base_url="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1",
api_key=os.environ["AWS_BEARER_TOKEN_BEDROCK"],
model_info=ModelInfo(vision=True, function_calling=True, json_output=True, family="unknown", structured_output=True),
)
messages = [UserMessage(content="Write a long poem about the sea.", source="user")]
args = {"max_completion_tokens": 5}
result = await client.create(messages, extra_create_args=args)
async for streamed in client.create_stream(messages, extra_create_args=args):
pass
print(f"{model}: create={result.finish_reason} create_stream={streamed.finish_reason}")
asyncio.run(main())
us.openai.gpt-6-sol: create=length create_stream=unknown
us.openai.gpt-6-luna: create=length create_stream=unknown
us.openai.gpt-6-astra: create=length create_stream=unknown
Expected behavior
create_stream reports the same finish_reason as create().
Additional context
Raw stream tail, proposed one-line fix, test RED/GREEN
Last two SSE events from us.openai.gpt-6-sol (trimmed; us.openai.gpt-6-astra has the same shape):
data: {"choices":[{"delta":{},"finish_reason":"stop","index":0}],...,"usage":{"completion_tokens":5,...}}
data: {"choices":[],...,"usage":{"completion_tokens":5,...}}
Fix: drop the chunk.usage is None condition. The stop_reason is None check already keeps an earlier finish reason when a LiteLLM-style usage chunk (with finish_reason=None) follows, and the existing usage-chunk tests still pass. The Azure AI client already reads finish_reason whenever it is set.
- stop_reason = choice.finish_reason if chunk.usage is None and stop_reason is None else stop_reason
+ stop_reason = choice.finish_reason if stop_reason is None else stop_reason
A mocked test (content chunk → finish chunk with usage → choices=[] usage chunk), parametrized for stop and length:
# main
E AssertionError: assert 'unknown' == 'stop'
E AssertionError: assert 'unknown' == 'length'
2 failed, 86 deselected in 2.68s
# with fix
2 passed, 86 deselected in 2.41s
tests/models/test_openai_model_client.py: 55 passed, 33 skipped
Astra has no tool-call row because Bedrock's chat completions endpoint returns 400 for function tools with Astra ("To use function tools, use /v1/responses"), on main and with the fix alike.
Released 0.7.5 has the same line. Happy to open a PR with the fix and the test.
Which packages was the bug in?
Python Extensions (autogen-ext)
AutoGen library version.
Python dev (main branch)
Other library version.
0.7.5 has the same code
Model used
us.openai.gpt-6-sol, us.openai.gpt-6-luna, us.openai.gpt-6-astra (and their global. profiles)
Model provider
AWS Bedrock
Other model provider
No response
Python version
3.11
.NET version
No response
Operating system
MacOS
What happened?
Thanks for the usage-chunk handling in
create_stream(#3972) and the normalized stop reasons (#5027). GPT-6 Sol, Luna and Astra on Amazon Bedrock already stream throughOpenAIChatCompletionClientwith justbase_url,api_keyandmodel_infoset. One result field still differs fromcreate().Describe the bug
create_streamreadsfinish_reasononly from chunks whereusage is None(_openai_client.pyL948). The Bedrock runtime endpoint/openai/v1/chat/completionsputs usage on the GPT-6 finish chunk, with or withoutinclude_usage. So the finish reason is never read and the streamed result says"unknown", including forlengthtruncation and tool calls. #4875 (DeepSeek,"ValueError: No stop reason found" always raised when llm usage returned) looks like the same code path; #5027 turned that error into"unknown".create()create_stream()todayus.openai.gpt-6-solandglobal.openai.gpt-6-sol: text / truncated / tool call¹us.openai.gpt-6-lunaandglobal.openai.gpt-6-luna: same 3create()us.openai.gpt-6-astra(us-east-1, us-west-2) andglobal.openai.gpt-6-astra(us-east-1): text / truncatedopenai.gpt-oss-120b-1:0(no usage on finish chunk)create()openai.gpt-6-sol(us-east-1, 3 cases) andopenai.gpt-6-astra(us-west-2, text / truncated), no usage on finish chunkcreate()¹ with
reasoning_effort="none", which Bedrock chat completions needs for Sol/Luna function tools.To Reproduce
Expected behavior
create_streamreports the samefinish_reasonascreate().Additional context
Raw stream tail, proposed one-line fix, test RED/GREEN
Last two SSE events from
us.openai.gpt-6-sol(trimmed;us.openai.gpt-6-astrahas the same shape):Fix: drop the
chunk.usage is Nonecondition. Thestop_reason is Nonecheck already keeps an earlier finish reason when a LiteLLM-style usage chunk (withfinish_reason=None) follows, and the existing usage-chunk tests still pass. The Azure AI client already readsfinish_reasonwhenever it is set.A mocked test (content chunk → finish chunk with usage →
choices=[]usage chunk), parametrized forstopandlength:Astra has no tool-call row because Bedrock's chat completions endpoint returns 400 for function tools with Astra ("To use function tools, use /v1/responses"), on main and with the fix alike.
Released 0.7.5 has the same line. Happy to open a PR with the fix and the test.
Which packages was the bug in?
Python Extensions (autogen-ext)
AutoGen library version.
Python dev (main branch)
Other library version.
0.7.5 has the same code
Model used
us.openai.gpt-6-sol, us.openai.gpt-6-luna, us.openai.gpt-6-astra (and their global. profiles)
Model provider
AWS Bedrock
Other model provider
No response
Python version
3.11
.NET version
No response
Operating system
MacOS