Files
DocsGPT/tests/e2e/fixtures/docs/readme.md
T
81b6ee5daa Pg 4 (#2390)
* feat: postgres tests

* feat: mongo cutoff

* feat: mongo cutoff

* feat: adjust docs and compose files

* fix: mini code mongo removals

* fix: tests and k8s mongo stuff

* feat: test fixes

* fix: ruff

* fix: vale

* Potential fix for pull request finding 'CodeQL / Clear-text logging of sensitive information'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* fix: mini suggestions

* vale lint fix 2

* fix: codeql columns thing

* fix: test mongo

* fix: tests coverage

* feat: better tests 4

* feat: more tests

* feat: decent coverage

* fix: ruff fixes

* fix: remove mongo mock

* feat: enhance workflow engine and API routes; add document retrieval and source handling

* feat: e2e tests

* fix: mcp, mongo and more

* fix: mini codeql warning

* fix: agent chunk view

* fix: mini issues

* fix: more pg fixes

* feat: postgres prep on start

* feat: qa tests

* fix: mini improvements

* fix: tests

---------

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
Co-authored-by: Siddhant Rai <siddhant.rai.5686@gmail.com>
2026-04-18 13:13:57 +01:00

1.9 KiB

Introduction to UTF-8 Encoding

UTF-8 is a variable-width character encoding capable of representing every character in the Unicode standard. It was designed as a backward-compatible replacement for ASCII, and it has become the dominant encoding for text on the web and in most modern file formats.

How UTF-8 Works

UTF-8 encodes each Unicode code point as one to four bytes. The number of bytes used depends on the numeric value of the code point. Characters in the ASCII range use a single byte identical to the ASCII byte, which is why any valid ASCII text is also a valid UTF-8 text. Higher code points use a lead byte that signals how many continuation bytes follow, and continuation bytes always begin with the bit pattern ten.

The design has a number of useful properties. Byte boundaries cannot be mistaken for character boundaries, because continuation bytes never look like lead bytes. A corrupted or truncated stream can be resynchronised by scanning forward to the next byte that is not a continuation byte. Sorting UTF-8 strings lexicographically by byte value produces the same order as sorting by Unicode code point.

Example

Below is a short Python snippet that encodes and decodes a UTF-8 string.

text = "hello"
data = text.encode("utf-8")
again = data.decode("utf-8")
assert again == text

Key Advantages

  • Compact for Latin-script text, because ASCII characters use only one byte.
  • Self-synchronising, which makes error recovery straightforward.
  • A strict superset of ASCII, so legacy ASCII tools handle it gracefully.
  • Well supported by every major programming language and operating system.
  • Avoids the byte-order ambiguity that affects UTF-16 and UTF-32.

UTF-8 is recommended by the Internet Engineering Task Force as the default encoding for web content, email headers, and most other text-based protocols. Using it consistently across an application removes an entire class of encoding-related bugs.