# UTF-8 Helpers

Single-character encoding, decoding, and validation helpers for UTF-8 text. Defined in [`./000 core.h`](https://facil.io/0.8.x/api/000-core/).

These functions work on one code point at a time and do not check memory bounds. Use them carefully — make sure the buffer has enough space for writes and is padded or NUL-terminated for reads.

## Compile-Time Flag

### `FIO_UTF8_ALLOW_IF`

Controls whether the UTF-8 helpers may use branchy (`if`) fast paths.

- **`1`** (default): ASCII-biased fast paths. Best when most input is ASCII.
- **`0`**: Branchless, constant-time-ish paths. Better when input distribution is unpredictable and branch mispredictions are a concern.

Defined before including [`./000 core.h`](https://facil.io/0.8.x/api/000-core/) to override the default.

## Length Helpers

### `fio_utf8_code_len`

```c
unsigned fio_utf8_code_len(uint32_t u);
```

Returns the number of bytes required to encode code point `u` in UTF-8.

Returns `1`–`4` for code points up to `0x1FFFFF`, and returns `0` only above that bound.

This bound is above the valid Unicode range (`0x10FFFF`), so the function accepts code points in the range `0x110000`–`0x1FFFFF` as a consequence of its bitmask. `fio_utf8_write` will encode those code points as 4-byte sequences.

### `fio_utf8_char_len_unsafe`

```c
unsigned fio_utf8_char_len_unsafe(uint8_t c);
```

Classifies a single leading byte without validating continuation bytes.

Returns:

- `1`–`4`: the expected total byte length of the UTF-8 character.
- `8`: a continuation byte (middle of a multi-byte character).
- `0`: an invalid leading byte.

Use this only to recover length information after a successful `fio_utf8_char_len`, `fio_utf8_write`, or similar call where validity is already known.

### `fio_utf8_char_len`

```c
unsigned fio_utf8_char_len(const void *str);
```

Returns the number of valid UTF-8 bytes used by the first character at `str`, validating continuation bytes.

Returns `0` if `str` does not point to a valid UTF-8-encoded code point. Returns `1`–`4` for a valid character.

## Write Helper

### `fio_utf8_write`

```c
unsigned fio_utf8_write(void *dest, uint32_t u);
```

Writes the UTF-8 encoding of code point `u` to `dest`. Returns the number of bytes written (`0`–`4`).

`u` is treated as described for `fio_utf8_code_len`: code points up to `0x1FFFFF` are encoded (including the non-Unicode range `0x110000`–`0x1FFFFF`), and higher values result in `0` bytes written.

The caller must ensure `dest` has room for up to 4 bytes. Typical use advances the destination pointer:

```c
dest += fio_utf8_write(dest, u);
```

## Read Helper

### `fio_utf8_read`

```c
uint32_t fio_utf8_read(char **str);
```

Decodes the first UTF-8 character at `*str` and returns its code point. Advances `*str` by the number of bytes consumed.

Returns `0` if the byte at `*str` is not a valid UTF-8 start byte or continuation sequence.

## Peek Helper

### `fio_utf8_peek`

```c
uint32_t fio_utf8_peek(const char *str);
```

Decodes the first UTF-8 character at `str` and returns its code point. Does not modify `str`.

Returns `0` for invalid input, same as `fio_utf8_read`.

## See Also

- [`./000 core.md`](https://facil.io/0.8.x/facil-io-stl-core/) — overview of core helpers.
- [`./000 core.h`](https://facil.io/0.8.x/api/000-core/) — source of truth for the UTF-8 helpers.
- [`./001 string info.md`](https://facil.io/0.8.x/string-and-buffer-information/) — string and buffer descriptors used with text helpers.
