Skip to content

use UTF-8 encoding in unit test test_output_to_file - #1494

Closed
oxygen dioxide (oxygen-dioxide) wants to merge 3 commits into
microsoft:mainfrom
oxygen-dioxide:encoding-fix
Closed

oxygen dioxide (oxygen-dioxide) wants to merge 3 commits into
microsoft:mainfrom
oxygen-dioxide:encoding-fix

Conversation

@oxygen-dioxide

@oxygen-dioxide oxygen dioxide (oxygen-dioxide) commented Dec 7, 2025 •

Copy link
Copy Markdown

On my Chinese windows, this test fails with the following error:

        with open(output_file, "r") as f:
>           output_data = f.read()
                          ^^^^^^^^
E           UnicodeDecodeError: 'gbk' codec can't decode byte 0xaa in position 61: illegal multibyte sequence

tests\test_cli_vectors.py:87: UnicodeDecodeError

because Markitdown writes to file in UTF-8, but the test tries to read it in local encoding

@afourney

Copy link
Copy Markdown
Member

Thanks for identifying the locale-dependent decoding failure in test_output_to_file. #2351 incorporated the same correction: the test now reads the generated output with an explicit UTF-8 encoding.

The change is on main, so I'm closing this PR as already addressed. Thank you for contributing the fix.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants