Repository navigation
Array-of-strings tuple sketch: key hashing is incompatible with Java (UTF-8 vs UTF-16) #533
Description
Activity
I have short time about two weeks.
What i think:
- We should change policy that using utf-8 as default in the AoS sketch. You can find last discussion. We need to discuss policy again.
- In c++, without third party dependencies, it is difficult to decode as utf-16LE(standard header is deprecated in c++17, removed in c++26). So as you reviewed in the previous merged PR, we need to add new dependency. It needs discussion too.
Because of this two reasons, i suggest revert my AoS sketch PR.
Releasing 5.3.0 is more important than introducing AoS sketch.
I will rewrite completely binary compatible AoS sketch.Two additions to include in the fix:
-
Hash input, precisely: the UTF-16 code units of the joined key, as little-endian bytes with no byte-order mark (UTF-16LE). This is what Java's
XxHash.hashCharArr(s.toCharArray(), ...)hashes. Note that a generic "UTF-16" converter is not the same thing: Java's ownStandardCharsets.UTF_16, for example, writes aFE FFbyte-order mark followed by big-endian bytes. -
Document the key-join behavior. Key elements are joined with
","without escaping, so different keys that join to the same string hash the same and are treated as one key, for example{"a,b"}and{"a", "b"}, or an empty array and{""}. We can't change this without breaking compatibility with Java, so all implementations must keep this behavior. Please note it in the API documentation, along these lines:Key arrays are hashed by joining their elements with
",". Keys that join to the same string, for example{"a,b"}and{"a", "b"}, are treated as the same key. If key elements may contain commas, escape or encode them before updating the sketch.
We'll add the same note to the Java documentation.
Reacted by Hyeonho Kim-
@proost ,
Your proposal makes sense. We should revert C++ AoS for 5.3.0. It removes the release blocker today, and the C++ AoS sketch never ships with incompatible hashing.I will open a disscussion on dev@ WRT the earlier UTF-8 validation discussion[1]. what we missed before was the already released Java AoS that used UTF-16, no BOM, which has been out for years. This discussion should not be rushed.
Claude helped me determine that the revert can be done cleanly: no later commits touch the 8 affected files, five can be deleted outright, and only two CMakeLists.txt files plus 39 lines in array_tuple_sketch.hpp need edits, with no downstream Python bindings at risk. Claude can test the revert locally without pushing to check if anything later depends on those 39 lines. If you are short on time I can have Claude do the revert for you. With your permission I can push the revert and open the PR for you to approve quickly, which will allow us to move forward on 5.3.0.
There are some ripple effects for the revert:
- Release notes: remove the "AoS tuple sketch (feat: AoS tuple sketch #476)" line from New Features.
- The TCK: the next C++ snapshot update will drop the aos_*_cpp.sk files. Java's cross-language test already skips missing files, so nothing breaks.
- Go Publish 2.2.0b1 under whylabs-datasketches name #191 stays open. Go has already released AoS, so it still needs fixing there, and the outcome of the dev@ discussion will decide how.
As for the rewrite, we don't need a 3rd party library to do a conversion from UTF-8 to UTF-16E: it is about 30 lines of code with no dependencies and reproduces Java's encoding exactly (what you also suggested in February). I could have Claude help with the rewrite also, but we need the result of the discussion for that and your go-ahead.
[1]: [DISCUSS] UTF-8 validation for string SerDe across sketches", 14 messages from Feb 14 to Mar 12, 2026.
https://lists.apache.org/thread/8p36zbmjp57fcjq7tlss69zy67459jv1The discussion on dev@ is open. https://lists.apache.org/thread/5vlbnodmw8h5s0oc1988f4pdhkjnv8qw
Thanks! Could you revert my merged PR?
No reverting feature in the mobile github app.
Sorry for late response.
- added a commit that references this issue
on Oct 3, 2026 Closing: the unreleased C++ array-of-strings sketch was reverted in #537, so 5.3.0 will not ship it with UTF-8 key hashing. The choice of hash encoding (UTF-16LE to match Java, or UTF-8 everywhere) is being discussed on dev@: https://lists.apache.org/thread/5vlbnodmw8h5s0oc1988f4pdhkjnv8qw. The AoS sketch will be rewritten once that is decided. The Go side is tracked in apache/datasketches-go#191.
@proost, while checking cross-language binary compatibility of the tuple sketches against the snapshots in datasketches-tck, we found that the C++ array-of-strings (AoS) tuple sketch hashes keys differently from Java. For the same keys, C++ and Java produce completely different hashes, so their sketches cannot be meaningfully combined.
Evidence
Comparing the TCK snapshots
aos_*_cpp.skandaos_*_java.sk, generated from the same keys:aos_1_n10aos_1_n1000aos_multikey_n1000aos_unicodeEach language reads the other's files without error, so nothing fails visibly. But a union of a Java sketch and a C++ sketch built from the same keys double-counts every key, and an intersection comes out empty.
Cause
The seed (
0x7A3CCA71) and the,separator match Java, but the bytes that get hashed don't:tuple/Util.stringArrHash) hashes the joined key as UTF-16 code units:XxHash.hashCharArrhashes eachcharas 2 bytes, little-endian.hash_array_of_strings_keyintuple/include/array_of_strings_sketch_impl.hpp) hashes the UTF-8 bytes of each string.Why C++ and Go should change, not Java
All three implementations need to agree. The cost of changing each one differs a lot:
So Java is the reference, and C++ and Go should match it.
Why this issue wasn't caught earlier
Our current cross-language tests only check that the .sk files can be read, not that the hashes match. We discovered this only recently, when we began Cross-Language Binary (CLB) testing.
Proposed fix
hash_array_of_strings_key, convert each key string from UTF-8 to UTF-16 code units, using surrogate pairs for code points above U+FFFF. Hash the code units as little-endian bytes, with,as2c 00between strings. Invalid UTF-8 should throwstd::invalid_argument.AosSketchCrossLanguageTest) and requires the hash sets to equal those in the Java.skfiles, includingaos_unicode.We have verified this approach: with UTF-8 to UTF-16LE conversion, C++ reproduces the hashes in every Java AoS snapshot exactly. That includes
aos_unicode, whose keys contain Korean, Cyrillic and emoji outside the BMP (🔑, 🗝️), and the multikey and 1,000,000-item cases. Apart from the hashing, the format already matches: every multi-entry Java AoS snapshot round-trips through C++ byte for byte.Optional, while in this code:
compact_array_of_strings_tuple_sketch::serialize()has no default serde, unlikedeserialize(). Calling it without passingdefault_array_of_strings_serde<>()compiles but fails at link time with an undefinedserde<array<std::string>>symbol.Timing
The C++ AoS sketch (#476) has not been released yet. We'd like to fix this before 5.3.0, which we plan to release very soon, so that the first release is compatible with Java. Otherwise, fixing it later would invalidate C++ users' saved sketches.
Could you take this on in the next few days? If you're short on time, let us know and we can prepare the PR for your review.
Go has the same issue; see apache/datasketches-go#191.