* feat: postgres tests * feat: mongo cutoff * feat: mongo cutoff * feat: adjust docs and compose files * fix: mini code mongo removals * fix: tests and k8s mongo stuff * feat: test fixes * fix: ruff * fix: vale * Potential fix for pull request finding 'CodeQL / Clear-text logging of sensitive information' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * fix: mini suggestions * vale lint fix 2 * fix: codeql columns thing * fix: test mongo * fix: tests coverage * feat: better tests 4 * feat: more tests * feat: decent coverage * fix: ruff fixes * fix: remove mongo mock * feat: enhance workflow engine and API routes; add document retrieval and source handling * feat: e2e tests * fix: mcp, mongo and more * fix: mini codeql warning * fix: agent chunk view * fix: mini issues * fix: more pg fixes * feat: postgres prep on start * feat: qa tests * fix: mini improvements * fix: tests --------- Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> Co-authored-by: Siddhant Rai <siddhant.rai.5686@gmail.com>
1.9 KiB
Introduction to UTF-8 Encoding
UTF-8 is a variable-width character encoding capable of representing every character in the Unicode standard. It was designed as a backward-compatible replacement for ASCII, and it has become the dominant encoding for text on the web and in most modern file formats.
How UTF-8 Works
UTF-8 encodes each Unicode code point as one to four bytes. The number of bytes used depends on the numeric value of the code point. Characters in the ASCII range use a single byte identical to the ASCII byte, which is why any valid ASCII text is also a valid UTF-8 text. Higher code points use a lead byte that signals how many continuation bytes follow, and continuation bytes always begin with the bit pattern ten.
The design has a number of useful properties. Byte boundaries cannot be mistaken for character boundaries, because continuation bytes never look like lead bytes. A corrupted or truncated stream can be resynchronised by scanning forward to the next byte that is not a continuation byte. Sorting UTF-8 strings lexicographically by byte value produces the same order as sorting by Unicode code point.
Example
Below is a short Python snippet that encodes and decodes a UTF-8 string.
text = "hello"
data = text.encode("utf-8")
again = data.decode("utf-8")
assert again == text
Key Advantages
- Compact for Latin-script text, because ASCII characters use only one byte.
- Self-synchronising, which makes error recovery straightforward.
- A strict superset of ASCII, so legacy ASCII tools handle it gracefully.
- Well supported by every major programming language and operating system.
- Avoids the byte-order ambiguity that affects UTF-16 and UTF-32.
UTF-8 is recommended by the Internet Engineering Task Force as the default encoding for web content, email headers, and most other text-based protocols. Using it consistently across an application removes an entire class of encoding-related bugs.