Skip to content

feat: add protocol-v2 multivariate flag evaluation - #744

Draft
roncohen wants to merge 10 commits into
mainfrom
feature/multivariate-evaluator-v2
Draft

roncohen wants to merge 10 commits into
mainfrom
feature/multivariate-evaluator-v2

Conversation

@roncohen

@roncohen roncohen commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Make evaluator v2 a clean major release: newEvaluator(flag) is the only evaluation entry point. Remove the v1 value-rule API and legacy serialization helpers.
  • Preserve generic variant values and use a discriminated resolved result, distinguishing selected undefined from failure and requiring success metadata.
  • Keep prepared Sets and percentage bounds private. Precompute lookups/bounds once, use a simple allocation scan, and preserve boolean cohort hashing.
  • Return structured diagnostics instead of console logging. Invalid numeric/date comparisons cannot match through negation.
  • Keep the unchanged Node SDK on published evaluator 1.1.1 via an alias. A root Yarn resolution prevents transparent workspace linking before Changesets bumps the evaluator to 2.0.0.

Validation

  • 215 evaluator tests and 231 Node SDK tests passed.
  • Evaluator typecheck, full workspace build, lint, and benchmarks passed.
  • Regression tests verify the runtime export surface, private prepared types, generic/discriminated results, cached Sets, numeric/date error handling, allocation boundaries, published-v1 cohorts, and dependency isolation.
  • Changesets release-plan validation selects evaluator 2.0.0 and a compatible Node SDK patch; unchanged SDKs are not migrated to the v2 protocol.

Performance

Prepared v2 membership checks remain roughly 0.7 microseconds per check with up to 100,000 candidates in local microbenchmarks. Percentage bounds are prepared once and scanned with a simple loop. Benchmarks are included in the package.

Scope

Still a draft; nothing published. SDK adoption of /flags and variant-valued public APIs follows separately.

@roncohen

Copy link
Copy Markdown
Contributor Author

Performance review fixes pushed:

  • Added newFlagEvaluator(flag) for repeated checks. SDKs should prepare on definition refresh and reuse the returned function; evaluateFlag is the one-shot convenience API.
  • ANY_OF and NOT_ANY_OF use prebuilt Sets, including under groups and negations. Prepared filters no longer retain the original candidate arrays. Scalar checks use hash lookups; array contexts do one lookup per examined context element.
  • Percentage validation and cumulative thresholds move out of the hot path. Allocation selection uses binary search.
  • Reuse per-rule diagnostic scratch storage and avoid result-copy/group-callback allocations. Returned diagnostics remain isolated between calls.
  • Fixed the inclusive maximum hash selecting a trailing 0% allocation; it now selects the last nonzero allocation.

Validation: 285 tests passed; typecheck, workspace build, and lint passed. Added repeatable Vitest benchmarks and tests that assert Set lookup usage and no repeated percentage reads.

Local benchmarks: prepared v2 membership checks stayed around 0.7 microseconds/check across 10, 1,000, and 100,000 candidates. At 100,000 candidates, the original PR implementation measured about 77 microseconds/check versus 0.68 microseconds prepared (roughly 113x faster, excluding preparation). Prepared v2 was also about 1.2x faster than prepared legacy in the included membership benchmark. Results are local microbenchmarks, not production throughput guarantees.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant