Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 10 additions & 10 deletions en/clice/design/compilation-context.md
Original file line number Diff line number Diff line change
Expand Up @@ -118,7 +118,7 @@ A name in either of the first two layers that no rule declares is reported as gu

Two extension requests drive the menu. **`clice/listConfigurations`** returns the declared tags with the active, selected and default names. **`clice/switchConfiguration`** validates the name and persists it as the selection; the running server keeps its configuration, and the request tells the client that a restart applies the choice. Switching is a restart on purpose: a server process holds one configuration's world — its loaded databases, dependency graph, contexts, artifacts and index — and a fresh start into that world is the cold start every session already performs, warm from the configuration's own cache. In VS Code the status bar shows the active configuration whenever the rules declare any; picking another one persists it and restarts the server, after which every open document goes through the ordinary open path and its diagnostics, inactive regions and semantic tokens flip. A restart in socket mode drops every connected client.

Each configuration owns an **index library** of its own under the cache store, `index/<tag>~<hash>` (`index/default` for a configuration without tags): the persisted symbol index, the compile-command snapshot it was built against, the context choices and the artifact records. Switching therefore never mixes rows of two configurations and never reindexes a configuration on the way back: the library it left is reopened as is, and the usual reconciliation picks up only what changed on disk meanwhile. Only the active configuration is indexed. PCH and PCM blobs stay in the shared content-addressed namespaces but carry the configuration in their keys, since the dependency stamps that vouch for a blob live in the configuration's library; header-context files are pure content and shared. Libraries of configurations no longer in use are not reclaimed automatically.
Each configuration owns an **index library** of its own under the cache store, `index/<tag>~<hash>` (`index/default` for a configuration without tags): the persisted symbol index, the compile-command snapshot it was built against, the context choices and the artifact records. Switching therefore never mixes rows of two configurations and never reindexes a configuration on the way back: the library it left is reopened as is, and the usual reconciliation picks up only what changed on disk meanwhile. Only the active configuration is indexed. PCH and PCM blobs stay in the shared content-addressed namespaces but carry the configuration in their keys, since the dependency records that vouch for a blob live in the configuration's library. Libraries of configurations no longer in use are not reclaimed automatically.

## Automatic Self-Containedness Detection

Expand All @@ -136,7 +136,7 @@ The header context is conceptually clear (host source file + include position),

> The prefix synthesis discussed here only applies to headers the user has opened. Headers on disk that are not open do not need special handling — they are included and processed normally when each source file is compiled. The indexing system collects symbol information from headers as part of each source file's indexing pass, and the header's per-file index shard merges the variants produced under different source files (see [Index Design](symbol-index.md) for the merging mechanism).

clice uses **prefix synthesis + `-include` injection**: based on the host source file and include position from the header context, it extracts all content before the target header along the include chain, synthesizes it into prefix code, writes it to a prefix file on disk, and then injects that file into the compilation command via Clang's `-include` flag. The core advantage of this approach is its natural compatibility with PCH optimization — `-include`'d files are processed through Clang's predefines buffer before the main file, so when the preamble PCH is built, the prefix content is baked into the **same** PCH as the header's own preamble region; even a header with no directives of its own (an X macro style `.def` file) gets a PCH built for its prefix. On PCH reuse Clang validates and subsumes matching `-include`s, so the prefix is never processed twice. PCHs are cached by content key (preamble text + canonical flags), so files with identical prefixes automatically share one. When the user subsequently edits the header body, the large volume of header inclusions in the prefix does not need to be reprocessed each time. Detailed rationale for this design choice is in the FAQ section below. PCH construction, caching, and invalidation mechanisms are described in [Incremental Compilation Design](incremental-parse.md).
clice uses **prefix synthesis + `-include` injection**: based on the host source file and include position from the header context, it extracts all content before the target header along the include chain, synthesizes it into prefix code held in memory, and then injects it into the compilation command via Clang's `-include` flag. The core advantage of this approach is its natural compatibility with PCH optimization — `-include`'d files are processed through Clang's predefines buffer before the main file, so when the preamble PCH is built, the prefix content is baked into the **same** PCH as the header's own preamble region; even a header with no directives of its own (an X macro style `.def` file) gets a PCH built for its prefix. On PCH reuse Clang validates and subsumes matching `-include`s, so the prefix is never processed twice. PCHs are cached by content key (preamble text + canonical flags), so files with identical prefixes automatically share one. When the user subsequently edits the header body, the large volume of header inclusions in the prefix does not need to be reprocessed each time. Detailed rationale for this design choice is in the FAQ section below. PCH construction, caching, and invalidation mechanisms are described in [Incremental Compilation Design](incremental-parse.md).

The synthesis process has four stages. The following example illustrates the process. Suppose the project has these files:

Expand Down Expand Up @@ -175,21 +175,21 @@ The user opens `math.h`, which is not in the CDB, so prefix code must be synthes

The content after `#include "utils.h"` in `main.cpp` (`int main() { ... }`) is truncated, and the content after `#include "math.h"` in `utils.h` (`void util_func();`) is also truncated. This prefix code restores the preprocessor state that `math.h` would see in the original compilation: `<vector>` and `<string>` have been included, and `DEBUG` is defined as 1.

The prefix code is written to a disk cache, named by the xxh3 hash of its content, so identical content is stored only once.
The prefix code is not written to disk: it is handed to the compiler together with the request, as a file that exists only in memory. It is split along the chain — the part cut from each file becomes its own in-memory file, placed in that file's directory and ending with an include of the next part — and each part is named by the xxh3 hash of its content, so identical content always gets the same name.

**Step 4: Inject into the compilation command.** Using the host `main.cpp`'s CDB compilation command, replace the source file path with `math.h` and inject the prefix file via the `-include` flag:
**Step 4: Inject into the compilation command.** Using the host `main.cpp`'s CDB compilation command, replace the source file path with `math.h` and inject the host's part of the prefix via the `-include` flag:

```bash
clang -std=c++17 -Iinclude -include cache/header_context/a1b2c3d4.h math.h
clang -std=c++17 -Iinclude -include .clice-a1b2c3d4.h math.h
```

Clang processes the `-include` prefix file first, then compiles `math.h`. The effect is that `math.h` is compiled as if at the `#include "math.h"` position in `main.cpp`, with the correct preprocessor state.

Several engineering details matter here. Include directives are resolved against the host command's real search paths and matched by **absolute path**, so same-named headers in different directories cannot be confused; since the prefix file lives in the cache directory, quoted includes with relative paths are rewritten to their resolved absolute paths, otherwise they would be looked up against the wrong base directory. Matching prefers includes outside `#if` blocks (an occurrence in an untaken branch must not shadow the real one); when the cut lands inside `#if` blocks (most commonly an include guard on an intermediate header), balancing `#endif`s are appended so the prefix stays well-formed — the guard condition is still evaluated by the compiler, preserving the semantics.
Several engineering details matter here. Include directives are resolved against the host command's real search paths and matched by **absolute path**, so same-named headers in different directories cannot be confused. Because each part of the prefix sits in the directory of the file it was cut from, every lookup inside it — a quoted include, a `__has_include` probe, an include spelled through a macro — resolves exactly as it does in that file. Matching prefers includes outside `#if` blocks (an occurrence in an untaken branch must not shadow the real one); when the cut lands inside `#if` blocks (most commonly an include guard on an intermediate header), balancing `#endif`s are appended so the prefix stays well-formed — the guard condition is still evaluated by the compiler, preserving the semantics.

Beyond the prefix there is also a **suffix**: the content after the include position (mirrored along the chain — the direct includer's remainder first, the host's last) is synthesized into a suffix file and injected by appending a single `#include` line to the header's buffer at compile time. X macro fragments embedded in enums or function bodies thus see their surrounding braces close — the token stream runs continuously through prefix, main file and suffix. When the cut lands inside `#if` blocks, the prefix closes them with `#endif`s and the suffix reopens the same depth with `#if 1`s, keeping both sides balanced. Includes of the header itself along the chain (other occurrences) are redirected to a disk snapshot of it — the header's own path is remapped to the buffer with the trailing suffix include, so keeping them verbatim would recurse forever. The appended line sits past the editor's visible content, and diagnostics inside the suffix file are never attributed to the header itself.
Beyond the prefix there is also a **suffix**: the content after the include position (mirrored along the chain — the direct includer's remainder first, the host's last) is synthesized into in-memory suffix files the same way and injected by appending a single `#include` line to the header's buffer at compile time. X macro fragments embedded in enums or function bodies thus see their surrounding braces close — the token stream runs continuously through prefix, main file and suffix. When the cut lands inside `#if` blocks, the prefix closes them with `#endif`s and the suffix reopens the same depth with `#if 1`s, keeping both sides balanced. Includes of the header itself along the chain (other occurrences) are redirected to a snapshot of its disk content — the header's own path is remapped to the buffer with the trailing suffix include, so keeping them verbatim would recurse forever. The appended line sits past the editor's visible content, and diagnostics inside the suffix are never attributed to the header itself. The synthesized files name nothing on disk: they are never reported as dependencies, and the symbols declared in them — and in the files they include — are left for the host's own indexing.

The synthesis result (host, prefix file path, suffix file path, content hash, and the include chain with a content snapshot) is owned by the editor's context state and outlives the editing session — closing and reopening the header reuses it. The chain files' content is embedded in the prefix and the compiler never opens them, so regular dependency tracking is blind to them — staleness is handled by the chain snapshot (two-layer mtime + content hash detection): any change to a chain file re-synthesizes the prefix. Saving a chain file forces content re-validation even when its mtime is unchanged.
The synthesis result (host, the synthesized files, and the include chain with the content versions it read) is owned by the editor's context state and outlives the editing session — closing and reopening the header reuses it. The chain files' content is embedded in the prefix and the compiler never opens them, so regular dependency tracking is blind to them — staleness is handled by the chain's recorded content versions: any change to a chain file discards the context, and the next compile synthesizes it again.

## Multi-Context in Indexing

Expand Down Expand Up @@ -218,9 +218,9 @@ If the index only records the result from one context, go-to-definition or find-

However, two GCC/Clang compiler extensions behave differently: `__INCLUDE_LEVEL__` is 0 under `-include` mode (the target file is the main file), whereas in the original compilation it reflects the nested include depth; `__BASE_FILE__` returns the target header's path under `-include` mode, whereas in the original compilation it returns the host source file's path. These two extensions are rarely used in practice and have no impact on language server functionality.

- **Why write the prefix file to disk instead of using a virtual file?**
- **Why keep the prefix in memory instead of writing it to disk?**

Clang supports a virtual file system, so in theory the synthesized prefix code could exist only in memory without being written to disk. The choice to write disk files is mainly for debuggability: when users encounter problems, they can directly inspect the prefix file contents on disk, making it easier to diagnose issues and file bug reports. There is no particularly deep technical reason.
A file on disk would be derived state with a life of its own: it needs a writable cache directory (without one, every header that needs its includer's context breaks), it must be kept in step with eviction, and it can surface to the user as if it were a source file. Placing it in the cache directory also moves relative lookups away from the file the text came from. In memory, the synthesized text exists exactly as long as the context it belongs to, and each part can be placed where its lookups resolve correctly.

- **Why not use an approach that "compiles the host source file as the main unit and stops at the target position"?**

Expand Down
14 changes: 7 additions & 7 deletions en/clice/design/incremental-parse.md
Original file line number Diff line number Diff line change
Expand Up @@ -64,12 +64,12 @@ An alternative is directly comparing content hashes: recompute the hash of every

clice combines both into a two-layer detection strategy, implemented once in the master's file table and shared by every consumer that needs file freshness (PCH validation, index staleness, disk polling):

- **Layer 1 (stat fast path)**: Compare each dependency's (size, mtime) stamp against the stamp recorded for the file version the PCH was built from. An equal stamp means the file is unchanged.
- **Layer 2 (content hash verification)**: For files whose stamp differs, recompute the xxh3 content hash and compare against the recorded one. If the content is in fact identical, the recorded stamp is repaired in place so the next check takes the fast path again; only a hash mismatch invalidates the PCH.
- **Layer 1 (stat fast path)**: The file table keeps one observation per file — the (size, mtime) a read saw and the hash of the bytes it read. When a dependency's current stat equals that observation, its content hash is known without reading.
- **Layer 2 (content hash verification)**: For files whose stat differs, read the file and hash it; the read becomes the file's new observation, so the next check takes the fast path again. The PCH is stale only when the current hash differs from the hash of the bytes it was built from.

Layer 1 filters out the vast majority of unchanged files (the common-case path). Layer 2 eliminates stamp false positives (build tool touches, VCS checkouts, etc.). The combined effect: PCH is rebuilt only when the content of a dependency file has actually changed.
Layer 1 filters out the vast majority of unchanged files (the common-case path). Layer 2 eliminates stat false positives (build tool touches, VCS checkouts, etc.). The combined effect: PCH is rebuilt only when the content of a dependency file has actually changed.

`DepsSnapshot` is the per-artifact record for two-layer detection, captured when a PCH build completes. It stores each dependency's file identity and observed version (plus markers for files that were missing at build time) and delegates the actual comparison to the shared file table.
`DepsSnapshot` is the per-artifact record for two-layer detection, captured when a PCH build completes. It stores each dependency's file identity and the content version the build consumed (plus markers for files that were missing at build time) and delegates the actual comparison to the shared file table. The stat observation belongs to the file, not to the artifact: one observation serves every artifact that depends on the file.

### Pull-Based Compilation

Expand Down Expand Up @@ -148,11 +148,11 @@ Multiple feature requests may simultaneously trigger a PCH build for the same co

### Dependency Snapshot Timing Guarantee

The `DepsSnapshot` build timestamp (`build_at`) is obtained **before** computing file hashes. This ordering ensures there is no time window in which a modification could be missed:
The content versions in a `DepsSnapshot` are hashes of the bytes the compiler actually consumed, taken from its own buffers, so a file modified during the build is never mistaken for what the build read: its new content fails the hash comparison on the next check. The build timestamp (`build_at`), obtained **before** the build starts, covers the one remaining case -- a dependency whose consumed bytes the worker could not hash:

If a file is modified during hash computation, its mtime will be later than `build_at`. On the next two-layer detection, Layer 1 will flag this file as "possibly modified," and Layer 2 will recompute its hash and discover the change.
If the file's mtime is no later than `build_at`, the disk still holds the bytes the build consumed, and their hash is taken from the file table. If it is later, the file may have changed during the build: the dependency is recorded without a version, reads as changed, and the next build captures it again.

If the order were reversed -- hashing first, then obtaining the timestamp -- a window could arise: a file modified between hash computation and timestamp acquisition would have an mtime no later than `build_at`, causing the modification to be missed.
If the order were reversed -- building first, then obtaining the timestamp -- a window could arise: a file modified during the build would have an mtime no later than `build_at`, and its new content would be recorded as the version the build consumed, causing the modification to be missed.

### Overall Compilation Flow

Expand Down
Loading
Loading