OdbDesignLib
OdbDesign ODB++ Parsing Library
 
Loading...
Searching...
No Matches
Utf8Sanitizer.h File Reference

UTF-8 validation and Windows-1252 to UTF-8 transcoding utilities. More...

#include <cstddef>
#include <string>
#include <string_view>
#include <google/protobuf/message.h>

Go to the source code of this file.

Functions

bool Odb::Lib::Text::IsValidUtf8 (const char *data, std::size_t size) noexcept
 Validates that a byte sequence is valid UTF-8 per RFC 3629.
 
bool Odb::Lib::Text::IsValidUtf8 (std::string_view s) noexcept
 Validates that a string_view contains valid UTF-8.
 
std::string Odb::Lib::Text::ToUtf8 (std::string_view input)
 Converts input to valid UTF-8.
 
void Odb::Lib::Text::SanitizeToUtf8 (std::string &s)
 In-place convenience overload.
 
void Odb::Lib::Text::AssertAllStringFieldsAreValidUtf8 (const google::protobuf::Message &msg, std::string_view msgName)
 Debug-only assertion to verify all string fields in a message are valid UTF-8.
 

Detailed Description

UTF-8 validation and Windows-1252 to UTF-8 transcoding utilities.

@why This exists because ODB++ design files exported from European EDA tools often contain legacy 8-bit text encoded in Windows-1252 (CP1252) or ISO-8859-1. When these raw bytes are copied verbatim into Protobuf string fields, they produce invalid UTF-8 on the wire. Strict Protobuf clients (including .NET's Google.Protobuf) refuse to parse such messages and throw InvalidProtocolBufferException: "String is invalid UTF-8".

This module provides:

  • RFC 3629 compliant UTF-8 validation
  • CP1252 → UTF-8 transcoding for invalid sequences
  • Zero-allocation fast path for already-valid UTF-8
See also
https://github.com/nam20485/OdbDesign/docs/grpc/server-utf8-sanitization-prompt.md

Definition in file Utf8Sanitizer.h.

Function Documentation

◆ AssertAllStringFieldsAreValidUtf8()

void Odb::Lib::Text::AssertAllStringFieldsAreValidUtf8 ( const google::protobuf::Message &  msg,
std::string_view  msgName 
)

Debug-only assertion to verify all string fields in a message are valid UTF-8.

Recursively walks the protobuf message via the reflection API and asserts that every string field contains valid UTF-8: singular strings, repeated strings, and strings nested inside sub-messages and map entries (map keys and values). bytes fields are intentionally skipped — they may hold arbitrary binary data. It is only active in debug builds (#ifndef NDEBUG); in release builds it compiles to a no-op.

On failure the offending field path and a hex-escaped sample of the value are written to stderr and the process aborts.

Parameters
msgThe protobuf message to validate
msgNameA descriptive name for the message type (used in assertion messages)

Definition at line 403 of file Utf8Sanitizer.cpp.

◆ IsValidUtf8() [1/2]

bool Odb::Lib::Text::IsValidUtf8 ( const char *  data,
std::size_t  size 
)
noexcept

Validates that a byte sequence is valid UTF-8 per RFC 3629.

Checks for:

  • Correct continuation byte ranges
  • No overlong encodings
  • No surrogate halves (U+D800..U+DFFF)
  • No code points above U+10FFFF
Parameters
dataPointer to the byte sequence
sizeLength in bytes
Returns
true if valid UTF-8, false otherwise

Definition at line 332 of file Utf8Sanitizer.cpp.

◆ IsValidUtf8() [2/2]

bool Odb::Lib::Text::IsValidUtf8 ( std::string_view  s)
noexcept

Validates that a string_view contains valid UTF-8.

Parameters
sThe string to validate
Returns
true if valid UTF-8, false otherwise

Definition at line 349 of file Utf8Sanitizer.cpp.

◆ SanitizeToUtf8()

void Odb::Lib::Text::SanitizeToUtf8 ( std::string &  s)

In-place convenience overload.

Replaces the contents of s with ToUtf8(s) only if validation fails. Avoids allocation when already valid UTF-8.

May throw std::bad_alloc when transcoding is required (see ToUtf8).

Parameters
sThe string to sanitize

Definition at line 395 of file Utf8Sanitizer.cpp.

◆ ToUtf8()

std::string Odb::Lib::Text::ToUtf8 ( std::string_view  input)

Converts input to valid UTF-8.

  • If input is already valid UTF-8, returns a copy unchanged.
  • Otherwise, repairs it: every maximal valid UTF-8 sequence is kept as-is and each remaining invalid byte is decoded as Windows-1252 and re-encoded as UTF-8. Transcoding the whole string as CP1252 would corrupt the already-valid parts of mixed input into mojibake.

CP1252 differs from ISO-8859-1 only in the 0x80–0x9F range. The five undefined CP1252 slots (0x81, 0x8D, 0x8F, 0x90, 0x9D) are mapped to U+FFFD (REPLACEMENT CHARACTER).

Never returns invalid UTF-8. May throw std::bad_alloc on allocation failure; callers in the serialization path handle it like any other allocation failure (the per-RPC try/catch translates it to INTERNAL).

Parameters
inputThe raw byte sequence from ODB++ files
Returns
A valid UTF-8 string

Definition at line 354 of file Utf8Sanitizer.cpp.