Skip to content

OpenAIChatCompletionClient.create_stream returns finish_reason "unknown" when the finish chunk also carries usage (seen with GPT-6 Sol/Luna/Astra on Amazon Bedrock) #8277

Description

@kimnamu

What happened?

Thanks for the usage-chunk handling in create_stream (#3972) and the normalized stop reasons (#5027). GPT-6 Sol, Luna and Astra on Amazon Bedrock already stream through OpenAIChatCompletionClient with just base_url, api_key and model_info set. One result field still differs from create().

Describe the bug

create_stream reads finish_reason only from chunks where usage is None (_openai_client.py L948). The Bedrock runtime endpoint /openai/v1/chat/completions puts usage on the GPT-6 finish chunk, with or without include_usage. So the finish reason is never read and the streamed result says "unknown", including for length truncation and tool calls. #4875 (DeepSeek, "ValueError: No stop reason found" always raised when llm usage returned) looks like the same code path; #5027 turned that error into "unknown".

live, main@027ecf0a create() create_stream() today with the fix below
runtime us-east-1 us.openai.gpt-6-sol and global.openai.gpt-6-sol: text / truncated / tool call¹ stop / length / function_calls ❌ unknown ×3 ✅ stop / length / function_calls
runtime us-east-1 us.openai.gpt-6-luna and global.openai.gpt-6-luna: same 3 same ❌ unknown ×3 ✅ same as create()
runtime us.openai.gpt-6-astra (us-east-1, us-west-2) and global.openai.gpt-6-astra (us-east-1): text / truncated stop / length ❌ unknown ×2 ✅ stop / length
runtime openai.gpt-oss-120b-1:0 (no usage on finish chunk) stop / length / function_calls ✅ same as create() ✅ unchanged
Mantle openai.gpt-6-sol (us-east-1, 3 cases) and openai.gpt-6-astra (us-west-2, text / truncated), no usage on finish chunk stop / length (/ function_calls) ✅ same as create() ✅ unchanged
streamed usage, GPT-6 Sol text — (8, 5) ✅ (8, 5) unchanged

¹ with reasoning_effort="none", which Bedrock chat completions needs for Sol/Luna function tools.

To Reproduce

import asyncio
import os

from autogen_core.models import ModelInfo, UserMessage
from autogen_ext.models.openai import OpenAIChatCompletionClient


async def main() -> None:
    for model in ["us.openai.gpt-6-sol", "us.openai.gpt-6-luna", "us.openai.gpt-6-astra"]:
        client = OpenAIChatCompletionClient(
            model=model,
            base_url="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1",
            api_key=os.environ["AWS_BEARER_TOKEN_BEDROCK"],
            model_info=ModelInfo(vision=True, function_calling=True, json_output=True, family="unknown", structured_output=True),
        )
        messages = [UserMessage(content="Write a long poem about the sea.", source="user")]
        args = {"max_completion_tokens": 5}
        result = await client.create(messages, extra_create_args=args)
        async for streamed in client.create_stream(messages, extra_create_args=args):
            pass
        print(f"{model}: create={result.finish_reason} create_stream={streamed.finish_reason}")


asyncio.run(main())
us.openai.gpt-6-sol: create=length create_stream=unknown
us.openai.gpt-6-luna: create=length create_stream=unknown
us.openai.gpt-6-astra: create=length create_stream=unknown

Expected behavior

create_stream reports the same finish_reason as create().

Additional context

Raw stream tail, proposed one-line fix, test RED/GREEN

Last two SSE events from us.openai.gpt-6-sol (trimmed; us.openai.gpt-6-astra has the same shape):

data: {"choices":[{"delta":{},"finish_reason":"stop","index":0}],...,"usage":{"completion_tokens":5,...}}
data: {"choices":[],...,"usage":{"completion_tokens":5,...}}

Fix: drop the chunk.usage is None condition. The stop_reason is None check already keeps an earlier finish reason when a LiteLLM-style usage chunk (with finish_reason=None) follows, and the existing usage-chunk tests still pass. The Azure AI client already reads finish_reason whenever it is set.

-            stop_reason = choice.finish_reason if chunk.usage is None and stop_reason is None else stop_reason
+            stop_reason = choice.finish_reason if stop_reason is None else stop_reason

A mocked test (content chunk → finish chunk with usage → choices=[] usage chunk), parametrized for stop and length:

# main
E       AssertionError: assert 'unknown' == 'stop'
E       AssertionError: assert 'unknown' == 'length'
2 failed, 86 deselected in 2.68s
# with fix
2 passed, 86 deselected in 2.41s
tests/models/test_openai_model_client.py: 55 passed, 33 skipped

Astra has no tool-call row because Bedrock's chat completions endpoint returns 400 for function tools with Astra ("To use function tools, use /v1/responses"), on main and with the fix alike.

Released 0.7.5 has the same line. Happy to open a PR with the fix and the test.

Which packages was the bug in?

Python Extensions (autogen-ext)

AutoGen library version.

Python dev (main branch)

Other library version.

0.7.5 has the same code

Model used

us.openai.gpt-6-sol, us.openai.gpt-6-luna, us.openai.gpt-6-astra (and their global. profiles)

Model provider

AWS Bedrock

Other model provider

No response

Python version

3.11

.NET version

No response

Operating system

MacOS

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions