diff --git a/en/clice/design/compilation-context.md b/en/clice/design/compilation-context.md index af722245..5911dd62 100644 --- a/en/clice/design/compilation-context.md +++ b/en/clice/design/compilation-context.md @@ -118,7 +118,7 @@ A name in either of the first two layers that no rule declares is reported as gu Two extension requests drive the menu. **`clice/listConfigurations`** returns the declared tags with the active, selected and default names. **`clice/switchConfiguration`** validates the name and persists it as the selection; the running server keeps its configuration, and the request tells the client that a restart applies the choice. Switching is a restart on purpose: a server process holds one configuration's world — its loaded databases, dependency graph, contexts, artifacts and index — and a fresh start into that world is the cold start every session already performs, warm from the configuration's own cache. In VS Code the status bar shows the active configuration whenever the rules declare any; picking another one persists it and restarts the server, after which every open document goes through the ordinary open path and its diagnostics, inactive regions and semantic tokens flip. A restart in socket mode drops every connected client. -Each configuration owns an **index library** of its own under the cache store, `index/~` (`index/default` for a configuration without tags): the persisted symbol index, the compile-command snapshot it was built against, the context choices and the artifact records. Switching therefore never mixes rows of two configurations and never reindexes a configuration on the way back: the library it left is reopened as is, and the usual reconciliation picks up only what changed on disk meanwhile. Only the active configuration is indexed. PCH and PCM blobs stay in the shared content-addressed namespaces but carry the configuration in their keys, since the dependency stamps that vouch for a blob live in the configuration's library; header-context files are pure content and shared. Libraries of configurations no longer in use are not reclaimed automatically. +Each configuration owns an **index library** of its own under the cache store, `index/~` (`index/default` for a configuration without tags): the persisted symbol index, the compile-command snapshot it was built against, the context choices and the artifact records. Switching therefore never mixes rows of two configurations and never reindexes a configuration on the way back: the library it left is reopened as is, and the usual reconciliation picks up only what changed on disk meanwhile. Only the active configuration is indexed. PCH and PCM blobs stay in the shared content-addressed namespaces but carry the configuration in their keys, since the dependency records that vouch for a blob live in the configuration's library. Libraries of configurations no longer in use are not reclaimed automatically. ## Automatic Self-Containedness Detection @@ -136,7 +136,7 @@ The header context is conceptually clear (host source file + include position), > The prefix synthesis discussed here only applies to headers the user has opened. Headers on disk that are not open do not need special handling — they are included and processed normally when each source file is compiled. The indexing system collects symbol information from headers as part of each source file's indexing pass, and the header's per-file index shard merges the variants produced under different source files (see [Index Design](symbol-index.md) for the merging mechanism). -clice uses **prefix synthesis + `-include` injection**: based on the host source file and include position from the header context, it extracts all content before the target header along the include chain, synthesizes it into prefix code, writes it to a prefix file on disk, and then injects that file into the compilation command via Clang's `-include` flag. The core advantage of this approach is its natural compatibility with PCH optimization — `-include`'d files are processed through Clang's predefines buffer before the main file, so when the preamble PCH is built, the prefix content is baked into the **same** PCH as the header's own preamble region; even a header with no directives of its own (an X macro style `.def` file) gets a PCH built for its prefix. On PCH reuse Clang validates and subsumes matching `-include`s, so the prefix is never processed twice. PCHs are cached by content key (preamble text + canonical flags), so files with identical prefixes automatically share one. When the user subsequently edits the header body, the large volume of header inclusions in the prefix does not need to be reprocessed each time. Detailed rationale for this design choice is in the FAQ section below. PCH construction, caching, and invalidation mechanisms are described in [Incremental Compilation Design](incremental-parse.md). +clice uses **prefix synthesis + `-include` injection**: based on the host source file and include position from the header context, it extracts all content before the target header along the include chain, synthesizes it into prefix code held in memory, and then injects it into the compilation command via Clang's `-include` flag. The core advantage of this approach is its natural compatibility with PCH optimization — `-include`'d files are processed through Clang's predefines buffer before the main file, so when the preamble PCH is built, the prefix content is baked into the **same** PCH as the header's own preamble region; even a header with no directives of its own (an X macro style `.def` file) gets a PCH built for its prefix. On PCH reuse Clang validates and subsumes matching `-include`s, so the prefix is never processed twice. PCHs are cached by content key (preamble text + canonical flags), so files with identical prefixes automatically share one. When the user subsequently edits the header body, the large volume of header inclusions in the prefix does not need to be reprocessed each time. Detailed rationale for this design choice is in the FAQ section below. PCH construction, caching, and invalidation mechanisms are described in [Incremental Compilation Design](incremental-parse.md). The synthesis process has four stages. The following example illustrates the process. Suppose the project has these files: @@ -175,21 +175,21 @@ The user opens `math.h`, which is not in the CDB, so prefix code must be synthes The content after `#include "utils.h"` in `main.cpp` (`int main() { ... }`) is truncated, and the content after `#include "math.h"` in `utils.h` (`void util_func();`) is also truncated. This prefix code restores the preprocessor state that `math.h` would see in the original compilation: `` and `` have been included, and `DEBUG` is defined as 1. -The prefix code is written to a disk cache, named by the xxh3 hash of its content, so identical content is stored only once. +The prefix code is not written to disk: it is handed to the compiler together with the request, as a file that exists only in memory. It is split along the chain — the part cut from each file becomes its own in-memory file, placed in that file's directory and ending with an include of the next part — and each part is named by the xxh3 hash of its content, so identical content always gets the same name. -**Step 4: Inject into the compilation command.** Using the host `main.cpp`'s CDB compilation command, replace the source file path with `math.h` and inject the prefix file via the `-include` flag: +**Step 4: Inject into the compilation command.** Using the host `main.cpp`'s CDB compilation command, replace the source file path with `math.h` and inject the host's part of the prefix via the `-include` flag: ```bash -clang -std=c++17 -Iinclude -include cache/header_context/a1b2c3d4.h math.h +clang -std=c++17 -Iinclude -include .clice-a1b2c3d4.h math.h ``` Clang processes the `-include` prefix file first, then compiles `math.h`. The effect is that `math.h` is compiled as if at the `#include "math.h"` position in `main.cpp`, with the correct preprocessor state. -Several engineering details matter here. Include directives are resolved against the host command's real search paths and matched by **absolute path**, so same-named headers in different directories cannot be confused; since the prefix file lives in the cache directory, quoted includes with relative paths are rewritten to their resolved absolute paths, otherwise they would be looked up against the wrong base directory. Matching prefers includes outside `#if` blocks (an occurrence in an untaken branch must not shadow the real one); when the cut lands inside `#if` blocks (most commonly an include guard on an intermediate header), balancing `#endif`s are appended so the prefix stays well-formed — the guard condition is still evaluated by the compiler, preserving the semantics. +Several engineering details matter here. Include directives are resolved against the host command's real search paths and matched by **absolute path**, so same-named headers in different directories cannot be confused. Because each part of the prefix sits in the directory of the file it was cut from, every lookup inside it — a quoted include, a `__has_include` probe, an include spelled through a macro — resolves exactly as it does in that file. Matching prefers includes outside `#if` blocks (an occurrence in an untaken branch must not shadow the real one); when the cut lands inside `#if` blocks (most commonly an include guard on an intermediate header), balancing `#endif`s are appended so the prefix stays well-formed — the guard condition is still evaluated by the compiler, preserving the semantics. -Beyond the prefix there is also a **suffix**: the content after the include position (mirrored along the chain — the direct includer's remainder first, the host's last) is synthesized into a suffix file and injected by appending a single `#include` line to the header's buffer at compile time. X macro fragments embedded in enums or function bodies thus see their surrounding braces close — the token stream runs continuously through prefix, main file and suffix. When the cut lands inside `#if` blocks, the prefix closes them with `#endif`s and the suffix reopens the same depth with `#if 1`s, keeping both sides balanced. Includes of the header itself along the chain (other occurrences) are redirected to a disk snapshot of it — the header's own path is remapped to the buffer with the trailing suffix include, so keeping them verbatim would recurse forever. The appended line sits past the editor's visible content, and diagnostics inside the suffix file are never attributed to the header itself. +Beyond the prefix there is also a **suffix**: the content after the include position (mirrored along the chain — the direct includer's remainder first, the host's last) is synthesized into in-memory suffix files the same way and injected by appending a single `#include` line to the header's buffer at compile time. X macro fragments embedded in enums or function bodies thus see their surrounding braces close — the token stream runs continuously through prefix, main file and suffix. When the cut lands inside `#if` blocks, the prefix closes them with `#endif`s and the suffix reopens the same depth with `#if 1`s, keeping both sides balanced. Includes of the header itself along the chain (other occurrences) are redirected to a snapshot of its disk content — the header's own path is remapped to the buffer with the trailing suffix include, so keeping them verbatim would recurse forever. The appended line sits past the editor's visible content, and diagnostics inside the suffix are never attributed to the header itself. The synthesized files name nothing on disk: they are never reported as dependencies, and the symbols declared in them — and in the files they include — are left for the host's own indexing. -The synthesis result (host, prefix file path, suffix file path, content hash, and the include chain with a content snapshot) is owned by the editor's context state and outlives the editing session — closing and reopening the header reuses it. The chain files' content is embedded in the prefix and the compiler never opens them, so regular dependency tracking is blind to them — staleness is handled by the chain snapshot (two-layer mtime + content hash detection): any change to a chain file re-synthesizes the prefix. Saving a chain file forces content re-validation even when its mtime is unchanged. +The synthesis result (host, the synthesized files, and the include chain with the content versions it read) is owned by the editor's context state and outlives the editing session — closing and reopening the header reuses it. The chain files' content is embedded in the prefix and the compiler never opens them, so regular dependency tracking is blind to them — staleness is handled by the chain's recorded content versions: any change to a chain file discards the context, and the next compile synthesizes it again. ## Multi-Context in Indexing @@ -218,9 +218,9 @@ If the index only records the result from one context, go-to-definition or find- However, two GCC/Clang compiler extensions behave differently: `__INCLUDE_LEVEL__` is 0 under `-include` mode (the target file is the main file), whereas in the original compilation it reflects the nested include depth; `__BASE_FILE__` returns the target header's path under `-include` mode, whereas in the original compilation it returns the host source file's path. These two extensions are rarely used in practice and have no impact on language server functionality. -- **Why write the prefix file to disk instead of using a virtual file?** +- **Why keep the prefix in memory instead of writing it to disk?** - Clang supports a virtual file system, so in theory the synthesized prefix code could exist only in memory without being written to disk. The choice to write disk files is mainly for debuggability: when users encounter problems, they can directly inspect the prefix file contents on disk, making it easier to diagnose issues and file bug reports. There is no particularly deep technical reason. + A file on disk would be derived state with a life of its own: it needs a writable cache directory (without one, every header that needs its includer's context breaks), it must be kept in step with eviction, and it can surface to the user as if it were a source file. Placing it in the cache directory also moves relative lookups away from the file the text came from. In memory, the synthesized text exists exactly as long as the context it belongs to, and each part can be placed where its lookups resolve correctly. - **Why not use an approach that "compiles the host source file as the main unit and stops at the target position"?** diff --git a/en/clice/design/incremental-parse.md b/en/clice/design/incremental-parse.md index ba476787..1c5b5d06 100644 --- a/en/clice/design/incremental-parse.md +++ b/en/clice/design/incremental-parse.md @@ -64,12 +64,12 @@ An alternative is directly comparing content hashes: recompute the hash of every clice combines both into a two-layer detection strategy, implemented once in the master's file table and shared by every consumer that needs file freshness (PCH validation, index staleness, disk polling): -- **Layer 1 (stat fast path)**: Compare each dependency's (size, mtime) stamp against the stamp recorded for the file version the PCH was built from. An equal stamp means the file is unchanged. -- **Layer 2 (content hash verification)**: For files whose stamp differs, recompute the xxh3 content hash and compare against the recorded one. If the content is in fact identical, the recorded stamp is repaired in place so the next check takes the fast path again; only a hash mismatch invalidates the PCH. +- **Layer 1 (stat fast path)**: The file table keeps one observation per file — the (size, mtime) a read saw and the hash of the bytes it read. When a dependency's current stat equals that observation, its content hash is known without reading. +- **Layer 2 (content hash verification)**: For files whose stat differs, read the file and hash it; the read becomes the file's new observation, so the next check takes the fast path again. The PCH is stale only when the current hash differs from the hash of the bytes it was built from. -Layer 1 filters out the vast majority of unchanged files (the common-case path). Layer 2 eliminates stamp false positives (build tool touches, VCS checkouts, etc.). The combined effect: PCH is rebuilt only when the content of a dependency file has actually changed. +Layer 1 filters out the vast majority of unchanged files (the common-case path). Layer 2 eliminates stat false positives (build tool touches, VCS checkouts, etc.). The combined effect: PCH is rebuilt only when the content of a dependency file has actually changed. -`DepsSnapshot` is the per-artifact record for two-layer detection, captured when a PCH build completes. It stores each dependency's file identity and observed version (plus markers for files that were missing at build time) and delegates the actual comparison to the shared file table. +`DepsSnapshot` is the per-artifact record for two-layer detection, captured when a PCH build completes. It stores each dependency's file identity and the content version the build consumed (plus markers for files that were missing at build time) and delegates the actual comparison to the shared file table. The stat observation belongs to the file, not to the artifact: one observation serves every artifact that depends on the file. ### Pull-Based Compilation @@ -148,11 +148,11 @@ Multiple feature requests may simultaneously trigger a PCH build for the same co ### Dependency Snapshot Timing Guarantee -The `DepsSnapshot` build timestamp (`build_at`) is obtained **before** computing file hashes. This ordering ensures there is no time window in which a modification could be missed: +The content versions in a `DepsSnapshot` are hashes of the bytes the compiler actually consumed, taken from its own buffers, so a file modified during the build is never mistaken for what the build read: its new content fails the hash comparison on the next check. The build timestamp (`build_at`), obtained **before** the build starts, covers the one remaining case -- a dependency whose consumed bytes the worker could not hash: -If a file is modified during hash computation, its mtime will be later than `build_at`. On the next two-layer detection, Layer 1 will flag this file as "possibly modified," and Layer 2 will recompute its hash and discover the change. +If the file's mtime is no later than `build_at`, the disk still holds the bytes the build consumed, and their hash is taken from the file table. If it is later, the file may have changed during the build: the dependency is recorded without a version, reads as changed, and the next build captures it again. -If the order were reversed -- hashing first, then obtaining the timestamp -- a window could arise: a file modified between hash computation and timestamp acquisition would have an mtime no later than `build_at`, causing the modification to be missed. +If the order were reversed -- building first, then obtaining the timestamp -- a window could arise: a file modified during the build would have an mtime no later than `build_at`, and its new content would be recorded as the version the build consumed, causing the modification to be missed. ### Overall Compilation Flow diff --git a/en/clice/design/overview.md b/en/clice/design/overview.md index 3f2ef1c5..72767ba4 100644 --- a/en/clice/design/overview.md +++ b/en/clice/design/overview.md @@ -27,7 +27,7 @@ General-purpose utilities and infrastructure shared by all other modules. ### `src/vfs/` — File Identity and Versions -- `FileTable`: Internalizes file paths as stable `Fid` identifiers used throughout the system, and owns the shared per-file facts derived from them — stat stamps, content versions, scan results, directory listings. The two-layer freshness check (stat fast path, then content hash with stamp repair) lives here once and is shared by every consumer: PCH validation, index staleness, and disk polling. +- `FileTable`: Internalizes file paths as stable `Fid` identifiers used throughout the system, and owns the shared per-file facts derived from them — the last observation of each file on disk, content versions, scan results, directory listings. The two-layer freshness check (stat fast path, then content hash) lives here once and is shared by every consumer: PCH validation, index staleness, and disk polling. Every look that finds other content than the last one is reported as a change, whoever looked. ### `src/config/` — Configuration @@ -126,9 +126,9 @@ The language server's core runtime, responsible for assembling all the layers ab - `Session` / `SessionStore`: The open-buffer truth for each open file — content, document version, generation, and serving state — created on didOpen and destroyed on didClose. Compile products do not live here - `ASTProjection` / `ASTProjectionTable`: The published products of each document's most recent compilation (feature results, PCH key, dependency snapshot) — an immutable read model replaced wholesale on each publication -- `EditorContext`: The editor's side of command resolution — the user's context choices, the header contexts resolved for open files, and the hosts of their synthesized preambles — layered over the project's `CommandResolver` for editor-facing compiles only -- `Invalidator`: The invalidation engine — folds file events (buffer opens/saves, on-disk changes, compilation-database reloads, worker crashes) into a deduplicated set of invalidation effects -- `FileTracker`: Stat-polling discovery of changes that happen outside the editor (a regenerated `compile_commands.json`, `git checkout`), feeding events to the `Invalidator`: the project's database watch plus a sweep of the files on disk, which judges each file against the content its include edges were scanned from +- `EditorContext`: The editor's side of command resolution — the user's context choices and the header contexts resolved for open files — layered over the project's `CommandResolver` for editor-facing compiles only +- `Invalidator`: The invalidation engine — folds file events (on-disk changes and removals, compilation-database reloads, worker crashes) into a deduplicated set of invalidation effects +- `FileTracker`: Stat-polling discovery of changes that happen outside the editor (a regenerated `compile_commands.json`, `git checkout`), feeding events to the `Invalidator`: the project's database watch plus a sweep, through the file table, of the files on disk and of the places a failed include lookup looked - `Quarantine`: Per-document crash accounting — documents whose content keeps killing workers are isolated and recover through licensed probe attempts **Services** — Read-side services consuming compilation and index results. diff --git a/en/clice/design/symbol-index.md b/en/clice/design/symbol-index.md index c48a035f..21a52a28 100644 --- a/en/clice/design/symbol-index.md +++ b/en/clice/design/symbol-index.md @@ -137,7 +137,7 @@ The index is rebuilt from the table, on the thread pool, when the symbols merged ### Staleness Detection -Staleness detection determines whether a file needs to be re-indexed. Each indexed artifact records the identity and observed content version of its inputs; validation is delegated to the master's shared file table, which performs the same two-layer check used everywhere else — a (size, mtime) stat fast path, then content-hash confirmation with stamp repair (see [Incremental Compilation](incremental-parse.md)). Re-indexing triggers only when input content actually changed; command changes are caught separately through the entry identity hashes recorded in the manifests. +Staleness detection determines whether a file needs to be re-indexed. Each indexed artifact records the identity and observed content version of its inputs; validation is delegated to the master's shared file table, which performs the same two-layer check used everywhere else — a (size, mtime) stat fast path against the file's last observation, then content-hash confirmation (see [Incremental Compilation](incremental-parse.md)). Re-indexing triggers only when input content actually changed; command changes are caught separately through the entry identity hashes recorded in the manifests. ### Background Indexing Scheduling diff --git a/zh/clice/design/compilation-context.md b/zh/clice/design/compilation-context.md index 2707db9d..03d7fbca 100644 --- a/zh/clice/design/compilation-context.md +++ b/zh/clice/design/compilation-context.md @@ -118,7 +118,7 @@ const char* error_message(int code) { 菜单由两个扩展请求驱动。**`clice/listConfigurations`** 返回已声明的标签,以及生效的、已选择的和默认的名称。**`clice/switchConfiguration`** 校验名称并将其持久化为选择;运行中的服务器仍保持原有配置,该请求会告知客户端重启后选择才会生效。切换要求重启是刻意的设计:一个服务器进程持有一个配置的整个世界——它加载的数据库、依赖图、上下文、产物和索引——重新启动进入那个世界,正是每次会话本来就要做的冷启动,而且能用上该配置自己的缓存。在 VS Code 中,只要规则声明了标签,状态栏就会显示生效的配置;选择另一个配置会将其持久化并重启服务器,之后每个打开的文档都会走一遍常规的打开流程,其诊断、非活跃区域和语义 Token 随之切换。在 socket 模式下,重启会断开所有已连接的客户端。 -每个配置在缓存存储下都拥有自己的**索引库** `index/~`(没有标签的配置为 `index/default`):持久化的符号索引、构建时所依据的编译命令快照、上下文选择和产物记录。因此,切换绝不会把两个配置的数据混在一起,切回去时也不会重新索引:离开时的索引库会原样重新打开,随后照常做一次核对,只处理这期间磁盘上发生变化的部分。只有生效的配置会被索引。PCH 与 PCM 产物仍留在共享的内容寻址命名空间中,但其键包含配置名,因为为产物背书的依赖戳记保存在该配置的索引库里;头文件上下文文件只取决于内容,仍然共享。不再使用的配置所占的索引库不会被自动回收。 +每个配置在缓存存储下都拥有自己的**索引库** `index/~`(没有标签的配置为 `index/default`):持久化的符号索引、构建时所依据的编译命令快照、上下文选择和产物记录。因此,切换绝不会把两个配置的数据混在一起,切回去时也不会重新索引:离开时的索引库会原样重新打开,随后照常做一次核对,只处理这期间磁盘上发生变化的部分。只有生效的配置会被索引。PCH 与 PCM 产物仍留在共享的内容寻址命名空间中,但其键包含配置名,因为为产物背书的依赖记录保存在该配置的索引库里。不再使用的配置所占的索引库不会被自动回收。 ## 自包含性自动检测 @@ -136,7 +136,7 @@ const char* error_message(int code) { > 这里讨论的前缀合成只适用于用户打开的头文件。磁盘上未打开的头文件不需要特殊处理——编译各个源文件时,它们会被正常包含和处理。索引系统在每个源文件的索引过程中收集头文件的符号信息,而头文件的逐文件索引分片会合并在不同源文件下产生的变体(合并机制见 [索引设计](symbol-index.md))。 -clice 采用**前缀合成 + `-include` 注入**方案:根据头文件上下文中的宿主源文件和 include 位置,沿 include 链提取目标头文件之前的所有内容,将其合成为前缀代码并写入磁盘上的前缀文件,然后通过 Clang 的 `-include` 标志将该文件注入编译命令。该方案的核心优势是天然兼容 PCH 优化——通过 `-include` 包含的文件会先经由 Clang 的 predefines 缓冲区处理,然后才处理主文件。因此,构建 Preamble PCH 时,前缀内容会与头文件自身的 Preamble 区域一起写入**同一个** PCH;即使头文件自身没有任何预处理指令(如 X macro 风格的 `.def` 文件),也会为其前缀构建 PCH。复用 PCH 时,Clang 会验证匹配的 `-include` 并将其视为已包含,因此前缀绝不会被处理两次。PCH 按内容键(Preamble 文本 + 规范化标志)缓存,前缀相同的文件会自动共享同一份 PCH。用户之后编辑头文件主体时,无需每次都重新处理前缀中包含的大量头文件。选择该设计的详细理由见下方 FAQ 章节。PCH 的构建、缓存和失效机制见 [增量编译设计](incremental-parse.md)。 +clice 采用**前缀合成 + `-include` 注入**方案:根据头文件上下文中的宿主源文件和 include 位置,沿 include 链提取目标头文件之前的所有内容,将其合成为保存在内存中的前缀代码,然后通过 Clang 的 `-include` 标志将其注入编译命令。该方案的核心优势是天然兼容 PCH 优化——通过 `-include` 包含的文件会先经由 Clang 的 predefines 缓冲区处理,然后才处理主文件。因此,构建 Preamble PCH 时,前缀内容会与头文件自身的 Preamble 区域一起写入**同一个** PCH;即使头文件自身没有任何预处理指令(如 X macro 风格的 `.def` 文件),也会为其前缀构建 PCH。复用 PCH 时,Clang 会验证匹配的 `-include` 并将其视为已包含,因此前缀绝不会被处理两次。PCH 按内容键(Preamble 文本 + 规范化标志)缓存,前缀相同的文件会自动共享同一份 PCH。用户之后编辑头文件主体时,无需每次都重新处理前缀中包含的大量头文件。选择该设计的详细理由见下方 FAQ 章节。PCH 的构建、缓存和失效机制见 [增量编译设计](incremental-parse.md)。 合成过程分为四个阶段。下面通过一个示例说明这一过程。假设项目中有以下文件: @@ -175,21 +175,21 @@ void util_func(); `main.cpp` 中 `#include "utils.h"` 之后的内容(`int main() { ... }`)会被截去,`utils.h` 中 `#include "math.h"` 之后的内容(`void util_func();`)也会被截去。这段前缀代码还原了 `math.h` 在原始编译中会看到的预处理器状态:`` 和 `` 已被包含,`DEBUG` 被定义为 1。 -前缀代码会写入磁盘缓存,并以其内容的 xxh3 哈希值命名,因此相同的内容只会存储一份。 +前缀代码不会写入磁盘:它随请求一起交给编译器,是一个只存在于内存中的文件。它沿包含链拆分——从每个文件截取的部分各自成为一个内存文件,放在该文件所在的目录中,并以包含下一部分的指令结尾——每一部分都以其内容的 xxh3 哈希值命名,因此相同的内容总会得到相同的名字。 -**第四步:注入编译命令。** 使用宿主 `main.cpp` 在 CDB 中的编译命令,将源文件路径替换为 `math.h`,并通过 `-include` 标志注入前缀文件: +**第四步:注入编译命令。** 使用宿主 `main.cpp` 在 CDB 中的编译命令,将源文件路径替换为 `math.h`,并通过 `-include` 标志注入前缀中属于宿主的那一部分: ```bash -clang -std=c++17 -Iinclude -include cache/header_context/a1b2c3d4.h math.h +clang -std=c++17 -Iinclude -include .clice-a1b2c3d4.h math.h ``` Clang 会先处理 `-include` 指定的前缀文件,然后编译 `math.h`。这样一来,`math.h` 就如同在 `main.cpp` 中 `#include "math.h"` 所在的位置进行编译,并具有正确的预处理器状态。 -这里有几个工程实现细节很重要。包含指令会根据宿主编译命令的实际搜索路径进行解析,并按**绝对路径**匹配,因此不会混淆不同目录下的同名头文件;由于前缀文件位于缓存目录中,其中使用引号指定相对路径的包含指令会被改写为解析后的绝对路径,否则查找时会采用错误的基准目录。匹配时优先选择 `#if` 块外的包含位置(未选中分支中的位置不得遮蔽实际生效的位置);当截断点位于 `#if` 块内时(最常见的情况是中间头文件中的头文件保护(include guard)),会追加用于配平的 `#endif`,使前缀保持良构——编译器仍会对保护条件求值,从而保持原有语义。 +这里有几个工程实现细节很重要。包含指令会根据宿主编译命令的实际搜索路径进行解析,并按**绝对路径**匹配,因此不会混淆不同目录下的同名头文件。由于前缀的每一部分都位于其截取来源文件所在的目录中,其中的每一次查找——引号形式的包含、`__has_include` 探测、通过宏拼写的包含——都会与在该文件中完全一样地解析。匹配时优先选择 `#if` 块外的包含位置(未选中分支中的位置不得遮蔽实际生效的位置);当截断点位于 `#if` 块内时(最常见的情况是中间头文件中的头文件保护(include guard)),会追加用于配平的 `#endif`,使前缀保持良构——编译器仍会对保护条件求值,从而保持原有语义。 -前缀之外还有**后缀**:包含位置之后的内容会沿包含链以镜像顺序拼接(直接包含者的剩余内容最先,宿主文件的剩余内容最后),合成为后缀文件;编译时,通过在头文件缓冲区末尾追加一行 `#include` 将其注入。因此,嵌入枚举或函数体中的 X macro 片段也能看到外围大括号闭合——Token 流连续穿过前缀、主文件和后缀。当截断点位于 `#if` 块内时,前缀会用 `#endif` 将其闭合,后缀则用相同层数的 `#if 1` 重新打开,使两侧保持平衡。包含链中对该头文件自身的其他包含位置会被重定向到它的磁盘快照——该头文件自身的路径已重映射到末尾附有后缀包含指令的缓冲区,因此若保持这些包含不变,就会无限递归。追加的这一行位于编辑器可见内容之外,后缀文件中的诊断绝不会归到该头文件本身。 +前缀之外还有**后缀**:包含位置之后的内容会沿包含链以镜像顺序拼接(直接包含者的剩余内容最先,宿主文件的剩余内容最后),以同样的方式合成为内存中的后缀文件;编译时,通过在头文件缓冲区末尾追加一行 `#include` 将其注入。因此,嵌入枚举或函数体中的 X macro 片段也能看到外围大括号闭合——Token 流连续穿过前缀、主文件和后缀。当截断点位于 `#if` 块内时,前缀会用 `#endif` 将其闭合,后缀则用相同层数的 `#if 1` 重新打开,使两侧保持平衡。包含链中对该头文件自身的其他包含位置会被重定向到其磁盘内容的快照——该头文件自身的路径已重映射到末尾附有后缀包含指令的缓冲区,因此若保持这些包含不变,就会无限递归。追加的这一行位于编辑器可见内容之外,后缀中的诊断绝不会归到该头文件本身。合成的文件不对应磁盘上的任何文件:它们绝不会被报告为依赖,其中声明的符号——以及它们所包含的文件中的符号——都留给宿主自身的索引处理。 -合成结果(宿主、前缀文件路径、后缀文件路径、内容哈希,以及附带内容快照的包含链)由编辑器的上下文状态持有,其生命周期超过编辑会话——关闭并重新打开头文件时会复用该结果。链上各文件的内容已嵌入前缀,编译器不会打开这些文件,因此常规依赖跟踪无法感知它们——包含链快照通过 mtime 与内容哈希两层检测来处理过期问题:链上任一文件发生变化,都会重新合成前缀。即使 mtime 未发生变化,保存链上文件也会强制重新验证其内容。 +合成结果(宿主、合成的文件,以及附带所读内容版本的包含链)由编辑器的上下文状态持有,其生命周期超过编辑会话——关闭并重新打开头文件时会复用该结果。链上各文件的内容已嵌入前缀,编译器不会打开这些文件,因此常规依赖跟踪无法感知它们——过期问题由包含链记录的内容版本处理:链上任一文件发生变化,都会丢弃该上下文,下一次编译时再重新合成。 ## 索引中的多上下文 @@ -218,9 +218,9 @@ Clang 会先处理 `-include` 指定的前缀文件,然后编译 `math.h`。 但有两个 GCC/Clang 编译器扩展的行为不同:`__INCLUDE_LEVEL__` 在 `-include` 模式下为 0(目标文件是主文件),而在原始编译中反映嵌套的 include 深度;`__BASE_FILE__` 在 `-include` 模式下返回目标头文件路径,而在原始编译中返回宿主源文件路径。这两个扩展在实际代码中极少使用,对语言服务器的功能没有影响。 -- **为什么将前缀文件写入磁盘,而不是使用虚拟文件?** +- **为什么将前缀保存在内存中,而不是写入磁盘?** - Clang 支持虚拟文件系统,因此理论上合成的前缀代码可以只存在于内存中,无须写入磁盘。选择写入磁盘主要是为了便于调试:用户遇到问题时,可以直接查看磁盘上的前缀文件内容,从而更容易诊断问题和提交缺陷报告。除此之外,没有什么深层的技术原因。 + 磁盘上的文件会成为一份有独立生命周期的派生状态:它需要一个可写的缓存目录(没有这样的目录,所有需要包含者上下文的头文件都会出错),需要与缓存淘汰保持同步,还可能像源文件一样出现在用户面前。把它放在缓存目录中,还会让相对路径的查找偏离文本的来源文件。放在内存中,合成文本的存续期与其所属的上下文完全一致,每一部分也都能放在让其中的查找正确解析的位置。 - **为什么不采用“将宿主源文件作为主编译单元,并在目标位置停止”的方案?** diff --git a/zh/clice/design/incremental-parse.md b/zh/clice/design/incremental-parse.md index c287b1b6..6caece7a 100644 --- a/zh/clice/design/incremental-parse.md +++ b/zh/clice/design/incremental-parse.md @@ -64,12 +64,12 @@ PCH 缓存 Preamble 中包含的所有头文件的预处理结果。只要任何 clice 将两者结合成两层检测策略,在 master 的文件表中统一实现,供所有需要判断文件是否为最新状态的使用方共享(PCH 校验、索引过期判断、磁盘轮询): -- **第一层(stat 快速路径)**:将每个依赖文件的 (size, mtime) 状态标记,与 PCH 构建所依据的该文件版本的记录值进行比较。状态标记相同即表示文件未变化。 -- **第二层(内容哈希校验)**:对于状态标记不同的文件,重新计算 xxh3 内容哈希并与记录值比较。如果内容实际上相同,就原地修正记录的状态标记,使下次检查再次采用快速路径;只有哈希不匹配才会使 PCH 失效。 +- **第一层(stat 快速路径)**:文件表为每个文件保存一份观察记录——读取时看到的 (size, mtime),以及所读字节的哈希。如果依赖文件当前的 stat 与该记录一致,无需读取就能得知其内容哈希。 +- **第二层(内容哈希校验)**:对于 stat 不同的文件,读取文件并计算哈希;这次读取成为该文件新的观察记录,使下次检查再次采用快速路径。只有当前哈希与构建 PCH 时所用字节的哈希不同,PCH 才算过期。 -第一层会过滤掉绝大多数未变化的文件(这是常见情况的处理路径)。第二层会消除状态标记造成的误报(构建工具更新时间戳、VCS 检出等)。二者结合后,只有依赖文件的内容实际发生变化,才会重建 PCH。 +第一层会过滤掉绝大多数未变化的文件(这是常见情况的处理路径)。第二层会消除 stat 造成的误报(构建工具更新时间戳、VCS 检出等)。二者结合后,只有依赖文件的内容实际发生变化,才会重建 PCH。 -`DepsSnapshot` 是用于两层检测、按产物保存的记录,在 PCH 构建完成时生成。它保存每个依赖文件的文件标识和当时观察到的版本(以及构建时缺失文件的标记),并将实际比较委托给共享文件表。 +`DepsSnapshot` 是用于两层检测、按产物保存的记录,在 PCH 构建完成时生成。它保存每个依赖文件的文件标识和构建所读取的内容版本(以及构建时缺失文件的标记),并将实际比较委托给共享文件表。stat 观察记录属于文件而不属于产物:依赖该文件的所有产物共用同一份观察记录。 ### 拉取式编译 @@ -148,11 +148,11 @@ PCH 构建由无状态工作进程执行(见[多进程架构](multi-process.md ### 依赖快照的时序保证 -`DepsSnapshot` 的构建时间戳(`build_at`)在计算文件哈希**之前**获取。这个顺序确保不存在遗漏修改的时间窗口: +`DepsSnapshot` 中的内容版本是编译器实际读取的字节的哈希,取自编译器自身的缓冲区,因此构建期间被修改的文件绝不会被误认为构建所读取的内容:它的新内容在下一次检查时无法通过哈希比较。构建时间戳(`build_at`)在构建开始**之前**获取,用于处理剩下的唯一一种情况——worker 无法为其所读字节计算哈希的依赖: -如果文件在哈希计算过程中被修改,其 mtime 会晚于 `build_at`。在下一次两层检测中,第一层会将该文件标记为“可能已修改”,第二层会重新计算其哈希并发现变化。 +如果文件的 mtime 不晚于 `build_at`,磁盘上仍是构建所读取的字节,其哈希直接取自文件表。如果晚于 `build_at`,该文件可能在构建期间发生了变化:这个依赖会被记录为没有版本,被视为已变化,并由下一次构建重新捕获。 -如果顺序反过来——先计算哈希,再获取时间戳——就可能出现一个时间窗口:文件在哈希计算与时间戳获取之间被修改时,其 mtime 不晚于 `build_at`,从而导致该修改被遗漏。 +如果顺序反过来——先构建,再获取时间戳——就可能出现一个时间窗口:构建期间被修改的文件,其 mtime 不晚于 `build_at`,它的新内容会被记录为构建所读取的版本,从而导致该修改被遗漏。 ### 整体编译流程 diff --git a/zh/clice/design/overview.md b/zh/clice/design/overview.md index f71e66db..e84c8ec5 100644 --- a/zh/clice/design/overview.md +++ b/zh/clice/design/overview.md @@ -27,7 +27,7 @@ clice 是一个全新的 C++ 语言服务器,从架构层面重新设计,旨 ### `src/vfs/` — 文件标识与版本 -- `FileTable`:将文件路径映射为全系统使用的稳定 `Fid` 标识,并管理据此得到的各文件共享信息——stat 标记、内容版本、扫描结果和目录列表。两层时效性检查(先走 stat 快速路径,再计算内容哈希并修复 stat 标记)集中实现在此处,供所有使用方共享:PCH 验证、索引过期判断和磁盘轮询。 +- `FileTable`:将文件路径映射为全系统使用的稳定 `Fid` 标识,并管理据此得到的各文件共享信息——每个文件在磁盘上最近一次的观察记录、内容版本、扫描结果和目录列表。两层时效性检查(先走 stat 快速路径,再计算内容哈希)集中实现在此处,供所有使用方共享:PCH 验证、索引过期判断和磁盘轮询。无论由谁查看,只要某次查看发现的内容与上一次不同,就会报告为一次变化。 ### `src/config/` — 配置 @@ -126,9 +126,9 @@ LSP 功能的具体实现。每个功能接收 `CompilationUnitRef`,返回对 - `Session` / `SessionStore`:记录每个打开文件的缓冲区权威状态——内容、文档版本、代次和服务状态——在 didOpen 时创建,在 didClose 时销毁。编译产物不存放在这里 - `ASTProjection` / `ASTProjectionTable`:每个文档最近一次编译所发布的产物(功能结果、PCH 键、依赖快照)——一种不可变的读取模型,每次发布时都会整体替换 -- `EditorContext`:命令解析中属于编辑器的一侧——用户的上下文选择、为打开文件解析出的头文件上下文,以及这些文件合成 Preamble 所用的宿主——叠加在项目的 `CommandResolver` 之上,仅用于面向编辑器的编译 -- `Invalidator`:失效引擎——归并文件事件(缓冲区打开/保存、磁盘内容变化、编译数据库重新加载、worker 崩溃),生成一组去重的失效操作 -- `FileTracker`:通过 stat 轮询发现编辑器之外发生的变化(重新生成 `compile_commands.json`、`git checkout`),并将事件送入 `Invalidator`:它包括对项目编译数据库的监视,以及对磁盘文件的一轮扫描——后者以扫描各文件 include 边时所依据的内容为基准,判断文件是否变化 +- `EditorContext`:命令解析中属于编辑器的一侧——用户的上下文选择,以及为打开文件解析出的头文件上下文——叠加在项目的 `CommandResolver` 之上,仅用于面向编辑器的编译 +- `Invalidator`:失效引擎——归并文件事件(磁盘内容变化与文件删除、编译数据库重新加载、worker 崩溃),生成一组去重的失效操作 +- `FileTracker`:通过 stat 轮询发现编辑器之外发生的变化(重新生成 `compile_commands.json`、`git checkout`),并将事件送入 `Invalidator`:它包括对项目编译数据库的监视,以及经由文件表进行的一轮扫描,覆盖磁盘上的文件和查找失败的 include 曾经查找过的位置 - `Quarantine`:按文档统计崩溃——内容屡次导致 worker 崩溃的文档会被隔离,并通过获准的探测尝试恢复 **服务** — 使用编译和索引结果的读取侧服务。 diff --git a/zh/clice/design/symbol-index.md b/zh/clice/design/symbol-index.md index 6c9abf0a..639574d5 100644 --- a/zh/clice/design/symbol-index.md +++ b/zh/clice/design/symbol-index.md @@ -137,7 +137,7 @@ Token 是模糊匹配在名称中可能走过的一条路径的 trigram,路径 ### 过期检测 -过期检测用于判断文件是否需要重新索引。每个索引产物都会记录其输入的标识以及观测到的内容版本;校验由主进程的共享文件表完成,它采用与其他各处相同的两层检查:先走 (size, mtime) stat 快速路径,再通过内容哈希确认,并修正状态标记(stamp)(见[增量编译](incremental-parse.md))。仅当输入内容确实发生变化时,才会触发重新索引;命令变更则通过清单中记录的条目标识哈希单独检测。 +过期检测用于判断文件是否需要重新索引。每个索引产物都会记录其输入的标识以及观测到的内容版本;校验由主进程的共享文件表完成,它采用与其他各处相同的两层检查:先以文件最近一次的观测记录为基准走 (size, mtime) stat 快速路径,再通过内容哈希确认(见[增量编译](incremental-parse.md))。仅当输入内容确实发生变化时,才会触发重新索引;命令变更则通过清单中记录的条目标识哈希单独检测。 ### 后台索引调度