UTF-8 validation and Windows-1252 to UTF-8 transcoding utilities. More...
#include <cstddef>#include <string>#include <string_view>#include <google/protobuf/message.h>Go to the source code of this file.
Functions | |
| bool | Odb::Lib::Text::IsValidUtf8 (const char *data, std::size_t size) noexcept |
| Validates that a byte sequence is valid UTF-8 per RFC 3629. | |
| bool | Odb::Lib::Text::IsValidUtf8 (std::string_view s) noexcept |
| Validates that a string_view contains valid UTF-8. | |
| std::string | Odb::Lib::Text::ToUtf8 (std::string_view input) |
| Converts input to valid UTF-8. | |
| void | Odb::Lib::Text::SanitizeToUtf8 (std::string &s) |
| In-place convenience overload. | |
| void | Odb::Lib::Text::AssertAllStringFieldsAreValidUtf8 (const google::protobuf::Message &msg, std::string_view msgName) |
| Debug-only assertion to verify all string fields in a message are valid UTF-8. | |
UTF-8 validation and Windows-1252 to UTF-8 transcoding utilities.
@why This exists because ODB++ design files exported from European EDA tools often contain legacy 8-bit text encoded in Windows-1252 (CP1252) or ISO-8859-1. When these raw bytes are copied verbatim into Protobuf string fields, they produce invalid UTF-8 on the wire. Strict Protobuf clients (including .NET's Google.Protobuf) refuse to parse such messages and throw InvalidProtocolBufferException: "String is invalid UTF-8".
This module provides:
Definition in file Utf8Sanitizer.h.
| void Odb::Lib::Text::AssertAllStringFieldsAreValidUtf8 | ( | const google::protobuf::Message & | msg, |
| std::string_view | msgName | ||
| ) |
Debug-only assertion to verify all string fields in a message are valid UTF-8.
Recursively walks the protobuf message via the reflection API and asserts that every string field contains valid UTF-8: singular strings, repeated strings, and strings nested inside sub-messages and map entries (map keys and values). bytes fields are intentionally skipped — they may hold arbitrary binary data. It is only active in debug builds (#ifndef NDEBUG); in release builds it compiles to a no-op.
On failure the offending field path and a hex-escaped sample of the value are written to stderr and the process aborts.
| msg | The protobuf message to validate |
| msgName | A descriptive name for the message type (used in assertion messages) |
Definition at line 403 of file Utf8Sanitizer.cpp.
|
noexcept |
Validates that a byte sequence is valid UTF-8 per RFC 3629.
Checks for:
| data | Pointer to the byte sequence |
| size | Length in bytes |
Definition at line 332 of file Utf8Sanitizer.cpp.
|
noexcept |
Validates that a string_view contains valid UTF-8.
| s | The string to validate |
Definition at line 349 of file Utf8Sanitizer.cpp.
| void Odb::Lib::Text::SanitizeToUtf8 | ( | std::string & | s | ) |
In-place convenience overload.
Replaces the contents of s with ToUtf8(s) only if validation fails. Avoids allocation when already valid UTF-8.
May throw std::bad_alloc when transcoding is required (see ToUtf8).
| s | The string to sanitize |
Definition at line 395 of file Utf8Sanitizer.cpp.
| std::string Odb::Lib::Text::ToUtf8 | ( | std::string_view | input | ) |
Converts input to valid UTF-8.
CP1252 differs from ISO-8859-1 only in the 0x80–0x9F range. The five undefined CP1252 slots (0x81, 0x8D, 0x8F, 0x90, 0x9D) are mapped to U+FFFD (REPLACEMENT CHARACTER).
Never returns invalid UTF-8. May throw std::bad_alloc on allocation failure; callers in the serialization path handle it like any other allocation failure (the per-RPC try/catch translates it to INTERNAL).
| input | The raw byte sequence from ODB++ files |
Definition at line 354 of file Utf8Sanitizer.cpp.