Implementation of UTF-8 validation and CP1252→UTF-8 transcoding. More...
#include "Utf8Sanitizer.h"#include <google/protobuf/descriptor.h>#include <google/protobuf/message.h>#include <cstdint>#include <cstdio>#include <cstdlib>#include <string>Go to the source code of this file.
Functions | |
| bool | Odb::Lib::Text::IsValidUtf8 (const char *data, std::size_t size) noexcept |
| Validates that a byte sequence is valid UTF-8 per RFC 3629. | |
| bool | Odb::Lib::Text::IsValidUtf8 (std::string_view s) noexcept |
| Validates that a string_view contains valid UTF-8. | |
| std::string | Odb::Lib::Text::ToUtf8 (std::string_view input) |
| Converts input to valid UTF-8. | |
| void | Odb::Lib::Text::SanitizeToUtf8 (std::string &s) |
| In-place convenience overload. | |
| void | Odb::Lib::Text::AssertAllStringFieldsAreValidUtf8 (const google::protobuf::Message &msg, std::string_view msgName) |
| Debug-only assertion to verify all string fields in a message are valid UTF-8. | |
Implementation of UTF-8 validation and CP1252→UTF-8 transcoding.
The CP1252→Unicode mapping table is derived from the official Microsoft mapping: https://www.unicode.org/Public/MAPPINGS/VENDORS/MICSFT/WINDOWS/CP1252.TXT
Undefined slots at 0x81, 0x8D, 0x8F, 0x90, 0x9D map to U+FFFD.
Definition in file Utf8Sanitizer.cpp.
| void Odb::Lib::Text::AssertAllStringFieldsAreValidUtf8 | ( | const google::protobuf::Message & | msg, |
| std::string_view | msgName | ||
| ) |
Debug-only assertion to verify all string fields in a message are valid UTF-8.
Recursively walks the protobuf message via the reflection API and asserts that every string field contains valid UTF-8: singular strings, repeated strings, and strings nested inside sub-messages and map entries (map keys and values). bytes fields are intentionally skipped — they may hold arbitrary binary data. It is only active in debug builds (#ifndef NDEBUG); in release builds it compiles to a no-op.
On failure the offending field path and a hex-escaped sample of the value are written to stderr and the process aborts.
| msg | The protobuf message to validate |
| msgName | A descriptive name for the message type (used in assertion messages) |
Definition at line 403 of file Utf8Sanitizer.cpp.
|
noexcept |
Validates that a byte sequence is valid UTF-8 per RFC 3629.
Checks for:
| data | Pointer to the byte sequence |
| size | Length in bytes |
Definition at line 332 of file Utf8Sanitizer.cpp.
|
noexcept |
Validates that a string_view contains valid UTF-8.
| s | The string to validate |
Definition at line 349 of file Utf8Sanitizer.cpp.
| void Odb::Lib::Text::SanitizeToUtf8 | ( | std::string & | s | ) |
In-place convenience overload.
Replaces the contents of s with ToUtf8(s) only if validation fails. Avoids allocation when already valid UTF-8.
May throw std::bad_alloc when transcoding is required (see ToUtf8).
| s | The string to sanitize |
Definition at line 395 of file Utf8Sanitizer.cpp.
| std::string Odb::Lib::Text::ToUtf8 | ( | std::string_view | input | ) |
Converts input to valid UTF-8.
CP1252 differs from ISO-8859-1 only in the 0x80–0x9F range. The five undefined CP1252 slots (0x81, 0x8D, 0x8F, 0x90, 0x9D) are mapped to U+FFFD (REPLACEMENT CHARACTER).
Never returns invalid UTF-8. May throw std::bad_alloc on allocation failure; callers in the serialization path handle it like any other allocation failure (the per-RPC try/catch translates it to INTERNAL).
| input | The raw byte sequence from ODB++ files |
Definition at line 354 of file Utf8Sanitizer.cpp.