diff --git a/llms.txt b/llms.txt index 09c0548..96e2f6d 100644 --- a/llms.txt +++ b/llms.txt @@ -29,6 +29,7 @@ Skills can be used with these AWS DevOps Agent types: - [VPC DNS Investigation Skill](skills/aws-vpc-dns-investigation/SKILL.md): Diagnoses VPC DNS resolution failures and validates DNS control-plane changes before they are applied, driving the aws-vpc-dns-diagnostics MCP server to observe live resolution from inside the affected subnet and to simulate a proposed change - [Bedrock Adoption Readiness Skill](skills/bedrock-adoption-readiness/SKILL.md): Assesses an AWS account's readiness to run Amazon Bedrock at production scale across IAM governance, data retention (ZDR), quota and capacity headroom, and operational observability, covering both the standard Bedrock and bedrock-mantle (OpenAI-compatible) surfaces with multi-region discovery - [Analytics OpenSearch Expertise Skill](skills/analytics-opensearch-expertise/SKILL.md): Performs read-only health assessments of Amazon OpenSearch Service domains through 24 deterministic checks across cluster health, storage and shards, performance, security, and cost optimization, producing a structured findings report with prioritized remediation guidance +- [Analytics DataOps Expertise Skill](skills/analytics-dataops-expertise/SKILL.md): Performs read-only DataOps maturity assessments of an AWS data platform through 26 questions scored 1-5 across five dimensions (architecture, security and governance, incident management and observability, automation and testing, and cost), producing a structured scorecard with per-dimension roll-ups and prioritized remediation guidance ## Key Concepts diff --git a/skills/analytics-dataops-expertise/.skilleval.yaml b/skills/analytics-dataops-expertise/.skilleval.yaml new file mode 100644 index 0000000..686a9c7 --- /dev/null +++ b/skills/analytics-dataops-expertise/.skilleval.yaml @@ -0,0 +1,3 @@ +audit: + ignore: + - STR-016 # README alongside SKILL.md is intentional diff --git a/skills/analytics-dataops-expertise/CHANGELOG.md b/skills/analytics-dataops-expertise/CHANGELOG.md new file mode 100644 index 0000000..25d96ca --- /dev/null +++ b/skills/analytics-dataops-expertise/CHANGELOG.md @@ -0,0 +1,42 @@ +# Changelog + +All notable changes to the `analytics-dataops-expertise` skill are documented here. + +## [1.0.2] - 2026-08-28 +### Fixed +- Dimension fidelity: pinned the scorecard to EXACTLY five dimensions (Architecture; + Security & Governance; Incident Management & Observability; Automation & Testing; Cost) + with fixed per-question membership, and aligned the frontmatter description to those five. + Fixes an observed run where the agent reported "7 dimensions" by promoting Architecture + questions (Real-time Processing Q4, Resilience Q7/Q8) into standalone top-level dimensions. + +## [1.0.1] - 2026-08-28 +### Fixed +- Report rendering: added a LITERAL-CONTENT RULE and an incremental RENDERING METHOD + to the Output Format so the agent writes every section and all 26 score-matrix rows + with real computed values. Fixes an observed failure in DevOps Agent where a saved + artifact contained placeholder labels ("full 3-tier content…", `{"q":"Q3"}` rows with + empty columns) instead of the actual report. Reinforced the same requirement in + Execution Flow step 6. + +## [1.0.0] - 2026-08-27 +### Added +- Initial release: 26 read-only control-plane maturity questions across five dimensions + (Architecture; Security & Governance; Incident Management & Observability; + Automation & Testing; Cost), each scored on a 1-5 maturity scale. +- Per-question checks that name the AWS API(s), the field(s) to read, and the signal-to-rating + mapping — 100% control-plane / API-driven with no data-plane access and no AWS-internal + data sources. +- Auto-score ceilings: questions that cannot be fully confirmed from control-plane signals + alone (Q6, Q8, Q11, Q13, Q14, Q19, Q20, Q23, Q24, Q26, Q30, Q37, Q38, Q39) are capped at 3; + the report states when a higher score requires conversational confirmation. +- Remediation Reference: a 26-entry verbatim dictionary (why-it-matters, resolution steps, + and verified official AWS documentation links) in `references/remediation-reference.md`, + loaded on demand via `read_skill_resource`. Every question scored ≤ 3 pulls its entry + verbatim; documentation links come only from the dictionary, eliminating URL hallucination. +- Structured scorecard output: per-section and overall averages, tiered findings, prioritized + recommendations, a raw-data reference, and a mandatory 26-row score matrix with a self-count + check. +- Error handling: missing services and region-unavailable APIs are treated as score-1 signals, + not blockers; a question is only SKIPPED when every API it needs is denied by IAM. Cost + Explorer must be enabled and is called in `us-east-1`. diff --git a/skills/analytics-dataops-expertise/README.md b/skills/analytics-dataops-expertise/README.md new file mode 100644 index 0000000..c7b6fb1 --- /dev/null +++ b/skills/analytics-dataops-expertise/README.md @@ -0,0 +1,62 @@ +# analytics-dataops-expertise + +A read-only **DataOps maturity assessment** skill for AWS DevOps Agent. Given an account ID and region, it runs control-plane API checks across the data platform, scores each dimension on a 1-5 maturity scale, and produces a structured scorecard with prioritized, doc-linked recommendations. + +## What it does + +Scores a customer's AWS data platform across five dimensions (26 maturity questions total): + +1. **Architecture** — catalog/governance patterns, real-time processing, change data capture, capacity planning, fault tolerance & HA, disaster recovery +2. **Security & Governance** — metadata management, lineage, data classification, data protection (encryption/PITR), fine-grained access control +3. **Incident Management & Observability** — workload analytics, monitoring & alerting, application tracing, drift detection, SLA management, user experience +4. **Automation & Testing** — infrastructure pipelines (IaC/CI-CD), orchestration, version/patch management, data quality testing +5. **Cost** — resource tagging, chargeback/showback, storage lifecycle management, data-processing cost, unused-resource cleanup + +Each question maps a set of read-only APIs to specific fields and a 1-5 rating rule. Questions that can't be fully confirmed from control-plane signals alone are **capped at 3** (existence of a resource ≠ mature use of it); the report states when a higher score needs conversational confirmation. + +It is **100% control-plane / API-driven** — no data-plane access (no catalog/table/index queries), no external packages, and no AWS-internal data sources. That makes it compatible with DevOps Agent without a custom MCP server. + +## Prerequisites + +- **Account ID and region** for the assessment — the scorer is account + region scoped. If either is missing, the skill asks before running. +- **AWS resources:** one or more analytics/data services in the target account/region (Glue, Kinesis, MSK, DMS, DynamoDB, RDS, OpenSearch, Redshift, EMR, MWAA, etc.). Absence of a service is itself a valid low-maturity signal. +- **Cost Explorer** must be enabled and is called in `us-east-1` (used by the two cost questions Q37/Q39). If it's unavailable, those score from tagging signals with the gap noted. + +**IAM permissions** the DevOps Agent execution role needs — all read-only. Most are covered by a ViewOnly/read-only managed policy; attach a supplemental policy for any gaps: + +- Glue / Lake Formation: `glue:GetDatabases`, `glue:GetJobs`, `glue:GetCrawlers`, `glue:ListRegistries`, `glue:ListSchemas`, `glue:ListDataQualityRulesets`, `glue:ListDataQualityResults`, `glue:ListWorkflows`, `lakeformation:GetDataLakeSettings`, `lakeformation:ListPermissions`, `lakeformation:ListDataCellsFilter` +- Streaming: `kinesis:ListStreams`, `firehose:ListDeliveryStreams`, `kinesisanalyticsv2:ListApplications`, `kafka:ListClustersV2` +- Migration/CDC: `dms:DescribeReplicationTasks`, `dms:DescribeReplicationInstances`, `dms:DescribeEndpoints`, `dms:DescribeEventSubscriptions` +- Databases: `dynamodb:ListTables`, `dynamodb:DescribeTable`, `dynamodb:DescribeContinuousBackups`, `rds:DescribeDBInstances`, `rds:DescribeDBEngineVersions`, `redshift:DescribeClusters`, `es:DescribeDomain`, `es:ListDomainNames` +- Scaling: `application-autoscaling:DescribeScalableTargets`, `autoscaling:DescribeAutoScalingGroups` +- Resilience: `backup:ListBackupPlans` +- Observability: `cloudwatch:DescribeAlarms`, `cloudwatch:ListDashboards`, `cloudwatch:GetMetricData`, `sns:ListTopics`, `events:ListRules`, `xray:GetSamplingRules`, `xray:GetGroups`, `rum:ListAppMonitors`, `quicksight:ListDashboards` +- Automation: `cloudformation:ListStacks`, `codepipeline:ListPipelines`, `codecommit:ListRepositories`, `mwaa:ListEnvironments`, `states:ListStateMachines`, `ssm:DescribePatchBaselines`, `lambda:ListFunctions` +- Governance/Security: `macie2:GetMacieSession`, `macie2:ListClassificationJobs`, `kms:ListKeys`, `s3:ListAllMyBuckets`, `s3:GetBucketLifecycleConfiguration`, `iam:ListPolicies`, `config:DescribeConfigRules`, `config:DescribeComplianceByConfigRule` +- Cost: `resourcegroupstaggingapi:GetResources`, `ce:GetCostAndUsage`, `ce:GetTags`, `cost-optimization-hub:ListRecommendations` + +## How to use it with DevOps Agent + +1. Get this skill directory onto your machine (clone the repo, or use GitHub's **Code → Download ZIP**), then zip just this skill folder: + ```bash + cd tools-for-devops-agent/skills + zip -r analytics-dataops-expertise.zip analytics-dataops-expertise + ``` + The zip must contain `SKILL.md` at the top of the `analytics-dataops-expertise/` folder. ZIP only, max 6 MB, no `scripts/` directory (this skill has none). +2. In the DevOps Agent Operator Web App, go to **Knowledge → Skills → Add skill → Upload skill** and select the zip. Choose the agent types that can use it (Generic applies to all; or target Chat / Incident RCA / etc.). Confirm it shows as active. +3. Prompt in natural language without naming the skill, and include the account ID and region, e.g.: + - "Assess the DataOps maturity of account `123456789012` in `eu-west-1`." + - "Review our data platform's architecture, governance, and cost posture. Account `123456789012`, region `us-east-1`." + - "How mature is our data pipeline observability and automation?" +4. Review the agent's reasoning trace to confirm the skill activated and the checks ran. Start a new chat after re-uploading a changed version so the latest skill loads. + +If the agent does not invoke the skill, refine the `description` field in `SKILL.md` (see "Optimizing description" in the Agent Skills specification). + +## Non-production disclaimer + +> ⚠️ This skill is sample code, not intended for production use without additional review and testing. Users should validate in a non-production environment first. + +## Maintainers + +- prasadnu +- genealpe diff --git a/skills/analytics-dataops-expertise/SKILL.md b/skills/analytics-dataops-expertise/SKILL.md new file mode 100644 index 0000000..d5b2f46 --- /dev/null +++ b/skills/analytics-dataops-expertise/SKILL.md @@ -0,0 +1,472 @@ +--- +name: analytics-dataops-expertise +description: "Amazon DataOps maturity assessment. Performs read-only, API-driven scoring of a customer's data platform across five fixed dimensions: Architecture, Security & Governance, Incident Management & Observability, Automation & Testing, and Cost. Activate this skill for requests about DataOps maturity, data-platform assessment, analytics maturity, data architecture review, data governance posture, pipeline/orchestration maturity, real-time/streaming data, or data cost optimization. Given an account ID and region, it scores 26 questions 1-5 from live account signals, rolls them up into those five dimensions, and produces a structured scorecard with prioritized recommendations. All checks use read-only AWS control-plane APIs (glue, kinesis, dms, rds, cloudwatch, kms, s3, iam, config, backup, mwaa, sfn, macie2, resourcegroupstaggingapi, costexplorer, costoptimizationhub) — no data-plane access required." +metadata: + version: "1.0.2" + author: prasadnu +--- + +# DataOps Maturity Assessment + +## Overview + +This skill performs a read-only DataOps maturity assessment of a customer's AWS +data platform using control-plane APIs available through the account's AWS +profile. It scores each dimension on a **1-5 maturity scale** and produces a +structured scorecard with findings and actionable recommendations. + +> **Scope:** CLI-agent compatible. Account-scoped, read-only, control-plane only. +> No external packages, no IDE workspace, no data-plane access (no queries +> against catalogs, tables, or indices). + +## Trigger Keywords + +`dataops`, `dataops maturity`, `data platform assessment`, `analytics maturity`, +`data architecture review`, `data governance`, `data platform health`, +`data maturity`, `data pipeline maturity`, `data cost optimization`, +`data observability` + +## Prerequisites + +- AWS CLI profile configured with read-only access to the target account +- **Account ID and region** for the assessment (the scorer is account + region scoped) +- IAM permissions required (all read-only). Most are covered by a read-only / + ViewOnly managed policy; attach the supplemental permissions for any gaps: + - `glue:GetDatabases`, `glue:GetJobs`, `glue:GetCrawlers`, `glue:ListRegistries`, `glue:ListSchemas`, `glue:ListDataQualityRulesets`, `glue:ListDataQualityResults`, `glue:ListWorkflows`, `glue:GetJob` + - `lakeformation:GetDataLakeSettings`, `lakeformation:ListPermissions`, `lakeformation:ListDataCellsFilter` + - `kinesis:ListStreams`, `firehose:ListDeliveryStreams`, `kinesisanalyticsv2:ListApplications`, `kafka:ListClustersV2` + - `dms:DescribeReplicationTasks`, `dms:DescribeReplicationInstances`, `dms:DescribeEndpoints`, `dms:DescribeEventSubscriptions` + - `dynamodb:ListTables`, `dynamodb:DescribeTable`, `dynamodb:DescribeContinuousBackups` + - `application-autoscaling:DescribeScalableTargets`, `autoscaling:DescribeAutoScalingGroups` + - `rds:DescribeDBInstances`, `rds:DescribeDBEngineVersions` + - `es:DescribeDomain`, `es:ListDomainNames`, `redshift:DescribeClusters` + - `backup:ListBackupPlans` + - `cloudwatch:DescribeAlarms`, `cloudwatch:ListDashboards`, `cloudwatch:GetMetricData` + - `sns:ListTopics`, `events:ListRules` + - `xray:GetSamplingRules`, `xray:GetGroups`, `rum:ListAppMonitors`, `quicksight:ListDashboards` + - `cloudformation:ListStacks`, `codepipeline:ListPipelines`, `codecommit:ListRepositories` + - `mwaa:ListEnvironments`, `states:ListStateMachines`, `ssm:DescribePatchBaselines` + - `lambda:ListFunctions` + - `macie2:GetMacieSession`, `macie2:ListClassificationJobs` + - `kms:ListKeys`, `s3:ListAllMyBuckets`, `s3:GetBucketLifecycleConfiguration` + - `iam:ListPolicies`, `config:DescribeConfigRules`, `config:DescribeComplianceByConfigRule` + - `resourcegroupstaggingapi:GetResources` + - `ce:GetCostAndUsage`, `ce:GetTags` (Cost Explorer must be enabled; call in `us-east-1`) + - `cost-optimization-hub:ListRecommendations` + +**AWS resources:** one or more analytics/data services in the target +account/region (Glue, Kinesis, MSK, DMS, DynamoDB, RDS, OpenSearch, Redshift, +EMR, MWAA, etc.). Absence of a service is itself a valid (low-maturity) signal. + +## Execution Flow + +### Input + +**Required:** account ID + region (e.g., `123456789012` / `eu-west-1`). + +If the account or region is missing, ask for it — do not assume. The assessment +runs against exactly the account + region provided. + +### Steps + +``` +1. Confirm target account ID + region with the user. +2. For each question, call its control-plane APIs (see Checks below). Any API + that errors or is not available in the region → treat as "not present" + (a score-1 signal), never as a blocker. +3. Score each question 1-5 using the rating scale + logic in its section. + Respect per-question auto-score ceilings (some questions cap at 3 — see below). +4. Roll up to per-section and overall averages. +5. Load the Remediation Reference — read_skill_resource(skill_id='analytics-dataops-expertise', + path='references/remediation-reference.md') — MANDATORY before writing recommendations. +6. Generate the scorecard with prioritized recommendations, following the Output + Format section EXACTLY. Write the full report content literally — the header, + Summary table, all findings, all recommendations, and all 26 matrix rows must + contain real computed values, never placeholder labels or a promise of content + (see the LITERAL-CONTENT RULE). If saving an artifact, verify it holds the same + fully-expanded content before finishing. +``` + +## Scoring model + +- Every question is scored **1 (lowest) to 5 (highest)** against the rating scale + in its section. Pick the highest rating whose conditions are fully met by the + observed signals; if signals fall between two ratings, choose the lower. +- **Auto-score ceiling (apply exactly).** Some questions cannot be fully + confirmed from control-plane signals alone (existence of a resource ≠ mature + use of it). These are capped — never score them above the cap from API signals; + state that ratings above the cap require conversational confirmation: + - **Cap = 3:** Q6, Q8, Q11, Q13, Q14, Q19, Q20, Q23, Q24, Q26, Q30, Q37, Q38, Q39 + - **No cap (score 1-5 from signals):** Q3, Q4, Q5, Q7, Q12, Q25, Q29, Q31, Q32, Q33, Q36, Q40 +- **API errors / missing services are score-1 signals**, not SKIPPED. A region + with no Kinesis/MSK/Flink is genuinely "no real-time processing" (Q4 = 1). Only + mark a question SKIPPED if EVERY API it needs is denied by IAM (state which). +- **No internal data.** This skill uses only account control-plane APIs. It does + NOT use support-case history, internal customer records, or any AWS-internal + data source. Do not infer or fabricate such signals. + +## Checks & Decision Logic + +Each check names the AWS API(s), the field(s) to read, and how observed signals +map to the 1-5 rating. All CloudWatch metrics use namespace as noted with the +account ID as the `ClientId`/account dimension where applicable. + +### Category 1: Architecture + +#### Q3 — Data Architecture Patterns (no cap) +- **APIs:** `glue.GetDatabases`, `glue.GetJobs`, `glue.GetCrawlers`, `lakeformation.GetDataLakeSettings` +- **Signals:** database/table count; active crawlers; Lake Formation governance (LF-Tags, non-`IAM_ALLOWED_PRINCIPALS` default perms); open-table-format indicators. +- **Rating:** + - 1: No catalog structure/governance; ad-hoc storage; no crawlers. + - 2: Basic catalog; few tables; no open formats; limited governance. + - 3: Structured catalog (≥3 databases, active crawlers); some governance; open formats emerging. + - 4: Well-organized catalog with clear domains; regular maintenance; comprehensive governance (LF-Tags in use). + - 5: Federated governance; automated maintenance + health alerting; self-serve discovery. + +#### Q4 — Real-time Data Processing (no cap) +- **APIs:** `kinesis.ListStreams`, `firehose.ListDeliveryStreams`, `kinesisanalyticsv2.ListApplications`, `kafka.ListClustersV2` +- **Signals:** count of streams / firehoses / running Flink apps / ACTIVE MSK clusters. +- **Rating:** + - 1: No streaming resources — all batch. + - 2: Basic streaming ingestion only (e.g., a stream or firehose), no stream processing. + - 3: Streaming ingestion + some real-time transforms (Flink/Kinesis Analytics present). + - 4: Comprehensive streaming pipelines with real-time transformations. + - 5: End-to-end streaming architecture with real-time analytics across critical data. + +#### Q5 — Change Data Capture (no cap) +- **APIs:** `dms.DescribeReplicationTasks/Instances/Endpoints/EventSubscriptions`, `dynamodb.ListTables` + `DescribeTable` (Streams) +- **Signals:** CDC-type DMS tasks (`cdc` / `full-load-and-cdc`); DMS event subscriptions; DynamoDB Streams enabled. +- **Rating:** + - 1: No CDC; full refreshes only. + - 2: Basic/manual CDC; limited validation; no monitoring. + - 3: Automated delta loads (DMS CDC tasks present); basic integrity validation. + - 4: Real-time CDC + automated integrity checks + monitoring/alerting (event subscriptions). + - 5: Intelligent change detection + automated recovery + cross-platform sync with audit trails. + +#### Q6 — Capacity Planning (cap 3) +- **APIs:** `application-autoscaling.DescribeScalableTargets` (dynamodb/ecs/kafka/lambda), `autoscaling.DescribeAutoScalingGroups` +- **Signals:** presence of auto-scaling policies; scheduled scaling. +- **Rating:** + - 1: No auto-scaling; purely manual. + - 2: Basic target tracking on some services; reactive. + - 3: Multiple services auto-scaled; some scheduled scaling. **(Max from signals.)** + - 4-5: Predictive/ML-based scaling + forecasting — requires conversation to confirm. + +#### Q7 — Fault Tolerance & High Availability (no cap) +- **APIs:** `rds.DescribeDBInstances` (MultiAZ, read replicas), `es.DescribeDomain` (ZoneAwareness), `kafka.ListClustersV2` (broker count/AZ), `redshift.DescribeClusters` +- **Signals:** Multi-AZ coverage across stateful services; read replicas; MSK multi-broker. +- **Rating:** + - 1: Single-AZ; no replicas/redundancy. + - 2: Multi-AZ on 1-2 services; no replicas. + - 3: Multi-AZ on primary data services; some read replicas. + - 4: Multi-AZ everywhere + automated failover + read replicas on critical DBs; MSK multi-broker. + - 5: All Rating 4 + cross-region replicas; full redundancy on all stateful services. + +#### Q8 — Disaster Recovery (cap 3) +- **APIs:** `backup.ListBackupPlans`, `rds.DescribeDBInstances` (PITR/`latestRestorableTime`) +- **Signals:** backup plans with recent `lastExecutionDate`; PITR coverage; cross-region backup copies. +- **Rating:** + - 1: No backup plans; no DR capability. + - 2: Backup plans exist; some PITR; no cross-region; no formal RPO/RTO. + - 3: Backup plans + cross-region/backup replication; long retention; PITR on critical DBs. **(Max from signals.)** + - 4-5: Defined/tested RPO-RTO, automated failover, DR drills — requires conversation to confirm. + +### Category 2: Security & Governance + +#### Q29 — Metadata Management & Catalog (no cap) +- **APIs:** `glue.GetDatabases`, `glue.GetCrawlers`, `lakeformation.GetDataLakeSettings` +- **Signals:** total tables; table description coverage; active crawlers; Lake Formation governed mode + LF-Tags. +- **Rating:** + - 1: No centralized catalog; manual discovery. + - 2: Basic/manual catalog (<50 tables, no crawlers); inconsistent metadata. + - 3: Automated metadata (≥50 tables OR active crawlers with ≥20 tables); standardized cataloging. + - 4: Comprehensive catalog + lineage + business context; integrated governance. + - 5: ML-powered classification/recommendations; self-maintaining metadata. + +#### Q30 — Lineage (cap 3) +- **APIs:** `glue.ListWorkflows`, `glue.GetJobs` (job bookmarks / dependencies) +- **Signals:** Glue Workflows (implicit lineage); job dependency graphs. Lineage is largely a data-plane/tooling property — do not over-credit from control-plane signals. +- **Rating:** + - 1: No lineage tracking. + - 2: Ad-hoc dependency mapping. + - 3: Lineage for select pipelines (Glue Workflows present), manual/inconsistent. **(Max from signals — 4-5 require confirmation of automated end-to-end lineage tooling.)** + +#### Q31 — Data Classification (no cap) +- **APIs:** `macie2.GetMacieSession`, `macie2.ListClassificationJobs`, `lakeformation.ListDataCellsFilter` +- **Signals:** Macie enabled + active classification jobs; LF cell-level filters. +- **Rating:** + - 1: No classification; Macie disabled. + - 2: Macie enabled but no active jobs. + - 3: Active Macie scanning + regular jobs; some LF cell filters; documented policy. + - 4: Comprehensive automated classification + LF-Tag-based policy enforcement. + - 5: ML-driven classification with predictive insights; automated remediation. + +#### Q32 — Data Protection (no cap) +- **APIs:** `kms.ListKeys`, `s3.ListAllMyBuckets` (+ per-bucket encryption), `rds.DescribeDBInstances` (`StorageEncrypted`, PITR), `dynamodb.DescribeContinuousBackups` +- **Signals:** customer-managed KMS keys; S3 default encryption; RDS storage encryption + PITR; DynamoDB PITR. +- **Rating:** + - 1: No/minimal encryption; no PITR. + - 2: Default encryption on some services; inconsistent. + - 3: Encryption at rest across most data services; PITR on critical stores. + - 4: CMK encryption + PITR everywhere + key rotation. + - 5: All Rating 4 + automated key management + cross-account/region protection controls. + +#### Q33 — Access Control (no cap) +- **APIs:** `lakeformation.ListPermissions`, `iam.ListPolicies`, `config.DescribeConfigRules`, `config.DescribeComplianceByConfigRule` +- **Signals:** Lake Formation permissions (fine-grained data access); least-privilege IAM; Config rules for governance + compliance status. +- **Rating:** + - 1: No fine-grained access control; broad IAM; no Config rules. + - 2: Basic IAM; no LF permissions; few/no Config rules. + - 3: LF permissions in use; Config rules present; some compliance tracking. + - 4: Comprehensive fine-grained access + Config compliance largely COMPLIANT. + - 5: Automated least-privilege + continuous compliance enforcement + attribute-based access. + +### Category 3: Incident Management & Observability + +#### Q11 — Workload Analytics (cap 3) +- **APIs:** `cloudwatch.ListDashboards` +- **Signals:** dashboards covering data-service namespaces; custom metrics. +- **Rating:** + - 1: No dashboards/custom metrics. + - 2: Some dashboards; default service metrics only. + - 3: Dedicated data dashboards; custom metrics; multiple data namespaces. **(Max — dashboard existence ≠ active KPI-driven ops; 4-5 need confirmation.)** + +#### Q12 — Monitoring & Optimization (no cap) +- **APIs:** `cloudwatch.DescribeAlarms`, `sns.ListTopics`, `events.ListRules` +- **Signals:** alarm coverage by data-service namespace; composite alarms; SNS topics; EventBridge rules (esp. → Lambda/SSM auto-remediation). +- **Rating:** + - 1: No alarms on data services. + - 2: Alarms on 1-2 data services; SNS linked; no EventBridge automation. + - 3: Alarms on ≥3 data services; SNS subscriptions; some EventBridge rules targeting data services. + - 4: Full alarm coverage across active data services; composite alarms; EventBridge → Lambda auto-remediation. + - 5: All Rating 4 + anomaly detectors on data metrics; confirmed auto-remediation; Contributor Insights. + +#### Q13 — Application Monitoring (cap 3) +- **APIs:** `xray.GetSamplingRules`, `xray.GetGroups`, `rum.ListAppMonitors` +- **Signals:** X-Ray sampling rules/groups; RUM monitors. +- **Rating:** + - 1: No X-Ray/RUM/tracing. + - 2: X-Ray on some services; default sampling. + - 3: Custom sampling + X-Ray groups; some RUM monitors. **(Max — 4-5 need ServiceLens/OTel confirmation.)** + +#### Q14 — Drift Detection (cap 3) +- **APIs:** `glue.ListRegistries`, `glue.ListSchemas`, `glue.ListDataQualityRulesets` +- **Signals:** Glue Schema Registry (schemas AVAILABLE); DQ rulesets. +- **Rating:** + - 1: No Schema Registry; no DQ rulesets. + - 2: One drift signal (a registry OR a ruleset). + - 3: Active schemas + DQ rulesets (multiple drift signals). **(Max — 4-5 need enforced-on-producers confirmation.)** + +#### Q19 — SLA Management (cap 3) +- **APIs:** `cloudwatch.DescribeAlarms` (treat-missing-data config, SLO-style alarms) +- **Signals:** alarms configured as SLIs/SLOs (latency/availability with sensible treat-missing-data). +- **Rating:** + - 1: No SLI/SLO alarms. + - 2: Basic availability alarms. + - 3: Multiple SLI/SLO-style alarms with proper treat-missing-data. **(Max — formal SLA program needs confirmation.)** + +#### Q20 — User Experience (cap 3) +- **APIs:** `rum.ListAppMonitors`, `quicksight.ListDashboards` +- **Signals:** RUM monitors; recently-published QuickSight dashboards. +- **Rating:** + - 1: No UX monitoring. + - 2: Some QuickSight dashboards; basic refresh visibility. + - 3: RUM + dashboards; multiple UX signals active. **(Max — staleness alerting/SLIs need confirmation.)** + +### Category 4: Automation & Testing + +#### Q23 — Infrastructure Pipeline (cap 3) +- **APIs:** `cloudformation.ListStacks`, `codepipeline.ListPipelines`, `codecommit.ListRepositories` +- **Signals:** active CFN stacks (IaC); CodePipeline pipelines; repos. +- **Rating:** + - 1: No IaC/pipelines; manual provisioning. + - 2: Some CFN stacks OR one pipeline. + - 3: IaC (active stacks) + CI/CD pipeline(s) present. **(Max — full GitOps/policy-as-code needs confirmation.)** + +#### Q24 — Orchestration (cap 3) +- **APIs:** `mwaa.ListEnvironments`, `states.ListStateMachines`, `glue.ListWorkflows` +- **Signals:** count of orchestration tools present (MWAA / Step Functions / Glue Workflows). +- **Rating:** + - 1: No orchestration; manual/cron. + - 2: One orchestration tool; basic scheduling. + - 3: Multiple orchestration tools; comprehensive coordination. **(Max — cross-pipeline SLA/self-healing needs confirmation.)** + +#### Q25 — Version Management (no cap) +- **APIs:** `rds.DescribeDBInstances` + `rds.DescribeDBEngineVersions`, `lambda.ListFunctions` (runtimes), `ssm.DescribePatchBaselines`, `es.DescribeDomain` (engine version) +- **Signals:** engine/runtime currency vs latest; deprecated Lambda runtimes; patch baselines. +- **Rating:** + - 1: No patch management; outdated versions; deprecated Lambda runtimes in use. + - 2: Some patched; 2+ minor version lag; no scheduled windows. + - 3: Most within 1-2 minor versions; some patch baselines; periodic upgrades. + - 4: Automated patching + within 1 minor version + scheduled windows. + - 5: All current; enforced patch baselines; zero deprecated runtimes; proactive cadence. + +#### Q26 — Data Quality Testing (cap 3) +- **APIs:** `glue.ListDataQualityRulesets`, `glue.ListDataQualityResults` +- **Signals:** DQ rulesets; recent DQ results (active execution). +- **Rating:** + - 1: No DQ rulesets. + - 2: Rulesets defined but no recent results. + - 3: Multiple rulesets + active execution (results present). **(Max — alerting-on-failure/coverage needs confirmation.)** + +### Category 5: Cost + +#### Q36 — Resource Tagging (no cap) +- **APIs:** `resourcegroupstaggingapi.GetResources` +- **Signals:** % of data resources tagged; presence of cost-allocation tags (Environment, Owner, CostCenter). +- **Rating:** + - 1: <30% tagged; no cost-allocation tags. + - 2: 30-60% tagged; ad-hoc tags. + - 3: 60-80% tagged; cost-allocation tags active; standard tags used. + - 4: >80% tagged; enforced tagging policy; periodic audits. + - 5: Near-100% tagged; automated tag enforcement (Config/SCP); tag-driven cost governance. + +#### Q37 — Chargeback / Showback (cap 3) +- **APIs:** `ce.GetCostAndUsage`, `ce.GetTags` *(Cost Explorer — enable it; call in `us-east-1`)* +- **Signals:** cost allocation tags active in CE; cost grouped by tag/service. +- **Rating:** + - 1: No cost allocation; no CE tag usage. + - 2: Some cost-allocation tags; basic CE reports; manual allocation. + - 3: Cost-allocation tags active + used for allocation across services. **(Max — automated chargeback needs confirmation.)** +- If CE is not enabled / AccessDenied → score from tagging signals and note the CE gap; do not SKIP unless no signal at all. + +#### Q38 — Storage Management (cap 3) +- **APIs:** `s3.ListAllMyBuckets` + `s3.GetBucketLifecycleConfiguration` +- **Signals:** buckets with lifecycle rules (tiering/expiration); Intelligent-Tiering. +- **Rating:** + - 1: No lifecycle rules; no tiering. + - 2: Lifecycle on some buckets; reactive cleanup. + - 3: Lifecycle/tiering on most data buckets; standard retention. **(Max — automated optimization needs confirmation.)** + +#### Q39 — Data Processing Cost (cap 3) +- **APIs:** `ce.GetCostAndUsage` (service filter), `emr.ListClusters`, `glue.GetJobs` +- **Signals:** EMR right-sizing/Spot; Glue worker sizing; processing cost trend from CE. +- **Rating:** + - 1: No processing-cost management; on-demand only. + - 2: Some cost awareness; manual reviews. + - 3: Cost allocation tags + some optimization (Spot/right-sizing evident). **(Max — ML-based/zero-waste needs confirmation.)** + +#### Q40 — Unused Resources (no cap) +- **APIs:** `cost-optimization-hub.ListRecommendations` +- **Signals:** count/value of idle/unused-resource recommendations; whether COH is active. +- **Rating:** + - 1: COH inactive / many unaddressed idle-resource recommendations; no cleanup cadence. + - 2: Manual periodic reviews; some Trusted Advisor checks; reactive cleanup. + - 3: COH active; recommendations generated; some cleanup cadence. + - 4: Regular recommendation review + remediation; low idle-resource footprint. + - 5: Automated detection + remediation; near-zero unused resources. + +### Remediation Reference rule (apply exactly) + +For EVERY question scored **≤ 3**, the Prioritized Recommendations section MUST +include that question's entry from `references/remediation-reference.md` (loaded +in Execution Flow step 5 via read_skill_resource): copy the "Why it matters" line +and "Resolve" steps verbatim, and include the "Dive deeper" link(s) EXACTLY as +written. NEVER emit a documentation link that is not present in the Remediation +Reference — do not construct, recall, or infer URLs from any other source. +Questions scored 4-5 get no remediation entry. + +## Output Format + +> **MANDATORY COVERAGE RULE:** The scorecard MUST evaluate and account for EVERY +> question (26 total: Q3, Q4, Q5, Q6, Q7, Q8, Q11, Q12, Q13, Q14, Q19, Q20, Q23, +> Q24, Q25, Q26, Q29, Q30, Q31, Q32, Q33, Q36, Q37, Q38, Q39, Q40). No question +> may be silently omitted. Before finishing, count the matrix rows: if not exactly +> 26, the report is incomplete — fix it before responding. +> **SECTION ROLL-UP:** each section's average = mean of its question scores. +> Overall = mean of all 26. Apply each question's rating scale and auto-score cap +> verbatim — never round a score up based on judgment. +> **EXACTLY FIVE DIMENSIONS (do not add, split, rename, or reorder).** The Summary +> table has EXACTLY these five sections with EXACTLY this question membership — this +> is fixed and must match the "Checks & Decision Logic" categories verbatim: +> - **Architecture** = Q3, Q4, Q5, Q6, Q7, Q8 +> - **Security & Governance** = Q29, Q30, Q31, Q32, Q33 +> - **Incident Mgmt & Observability** = Q11, Q12, Q13, Q14, Q19, Q20 +> - **Automation & Testing** = Q23, Q24, Q25, Q26 +> - **Cost** = Q36, Q37, Q38, Q39, Q40 +> +> Real-time processing (Q4), CDC (Q5), capacity planning (Q6), fault tolerance (Q7), +> and disaster recovery (Q8) are QUESTIONS INSIDE the Architecture dimension — they +> are NEVER promoted to their own top-level dimensions. Reporting "7 dimensions", a +> "Real-time Processing" dimension, or a "Resilience" dimension is a SPEC VIOLATION — +> the scorecard must show five dimensions and five only. +> +> **LITERAL-CONTENT RULE (apply exactly — no placeholders).** Every section and +> every table row MUST contain the actual computed content, fully written out. +> This is the single most important rule for the report: +> - NEVER write a description, label, summary, or promise of content in place of +> the content itself. Text like "full 3-tier content", "…included as persisted", +> "see full artifact", "25 recommendations here", or a matrix row of the form +> `{"q":"Q3"}` with empty columns is a FAILED report — it must be regenerated. +> - The Score Matrix MUST have 26 fully-populated rows: every row states the actual +> Dimension, the numeric Score, the Cap, the real Observed signal (actual counts / +> config values from the API calls), and the Rating rule applied. No empty cells. +> - Detailed Findings MUST list each question by name with its real observed signal +> and score. Prioritized Recommendations MUST contain the actual verbatim +> "Why it matters" + "Resolve" + "Dive deeper" text (see rule above) for every +> question scored ≤ 3 — not a count of them. +> - The report's FIRST content MUST be the `# DataOps Maturity Assessment` header +> and the Summary table, rendered literally with the real account ID, region, +> timestamp, and scores — not a description of the header. +> +> **RENDERING METHOD (avoid truncation on large output).** This report is long +> (26 questions + up to 25 remediation entries). Do NOT try to emit it in one +> oversized write that gets summarized or truncated. Build it incrementally and +> in full: write the header + Summary table first, then append Detailed Findings, +> then append each Prioritized Recommendation, then append the 26 matrix rows. +> If you save the report as an artifact, the artifact MUST contain this same fully +> expanded content — re-open/verify it after writing and, if any section shows a +> placeholder label or the matrix has fewer than 26 populated rows, rewrite it with +> the real content before telling the user it is done. + +Structure the report as: + +```markdown +# DataOps Maturity Assessment +## Account: {account_id} | Region: {region} | Assessed: {timestamp} + +## Summary +| Section | Avg (1-5) | Status | +|---------|-----------|--------| +| Architecture | {avg} | 🟢/🟡/🔴 | +| Security & Governance | {avg} | 🟢/🟡/🔴 | +| Incident Mgmt & Observability | {avg} | 🟢/🟡/🔴 | +| Automation & Testing | {avg} | 🟢/🟡/🔴 | +| Cost | {avg} | 🟢/🟡/🔴 | +| **Overall** | **{avg}** | | + +Status heuristic: avg < 2 → 🔴, 2-3.5 → 🟡, > 3.5 → 🟢. + +## Detailed Findings +### 🔴 Lowest maturity (act now — scored 1) +{questions} +### 🟡 Developing (scored 2-3) +{questions} +### 🟢 Mature (scored 4-5) +{questions} + +## Recommendations (prioritized) +1. {highest-impact, lowest-score first} +2. ... + +## Raw Data Reference +{key config values and counts used for scoring, per question} + +## Score Matrix (REQUIRED — one row per question, all 26, in order) +| Q | Dimension | Score (1-5) | Cap | Observed signal | Rating rule applied | +|---|-----------|-------------|-----|-----------------|---------------------| +| Q3 | Data Architecture | {n} | — | {value} | {rule} | +| ... (every question through Q40 — SKIPPED rows must state the IAM reason) | +``` + +## Error Handling + +- If an API returns `AccessDeniedException` → if the question has other signals, + score from those and note the gap; only mark a question SKIPPED if EVERY API it + needs is denied (state which permission is missing). +- If an API is not available in the region or returns empty → treat as "resource + not present" (a score-1 signal), not an error. +- Cost Explorer (`ce.*`) must be enabled and is called in `us-east-1`; if it is + unavailable, score Q37/Q39 from tagging/other signals and note the CE gap. +- This skill never uses AWS-internal data sources; if a signal cannot be obtained + from account control-plane APIs, say so — do not fabricate it. diff --git a/skills/analytics-dataops-expertise/evals/benchmark.json b/skills/analytics-dataops-expertise/evals/benchmark.json new file mode 100644 index 0000000..7a879a1 --- /dev/null +++ b/skills/analytics-dataops-expertise/evals/benchmark.json @@ -0,0 +1,834 @@ +{ + "skill_name": "analytics-dataops-expertise", + "skill_path": "/Users/prasadnu/Dante/OpenSearch-OpsReview-KiroPower/external-repos/tools-for-devops-agent/skills/analytics-dataops-expertise", + "eval_count": 8, + "runs_per_eval": 1, + "metadata": { + "timestamp": "2026-08-28T11:35:54Z" + }, + "runs": [ + { + "eval_id": "dataops-five-dimensions", + "run_index": 0, + "with_skill": { + "eval_id": "dataops-five-dimensions", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'Architecture' or contains 'architecture'", + "passed": true, + "evidence": "OR satisfied by: contains 'architecture' \u2014 Substring found: 'architecture'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Security' or contains 'security'", + "passed": true, + "evidence": "OR satisfied by: contains 'security' \u2014 Substring found: 'security'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Observability' or contains 'observability'", + "passed": true, + "evidence": "OR satisfied by: contains 'observability' \u2014 Substring found: 'observability'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Automation' or contains 'automation'", + "passed": true, + "evidence": "OR satisfied by: contains 'automation' \u2014 Substring found: 'automation'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Cost' or contains 'cost'", + "passed": true, + "evidence": "OR satisfied by: contains 'cost' \u2014 Substring found: 'cost'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 185, + "cache_read_input_tokens": 11306, + "cache_creation_input_tokens": 29171 + } + }, + "timing": { + "elapsed_seconds": 11.746247624978423 + }, + "raw_output": "According to the skill, the five dimensions (sections) scored are:\n\n1. **Architecture**\n2. **Security & Governance**\n3. **Incident Management & Observability**\n4. **Automation & Testing**\n5. **Cost**\n\nEach is scored on a 1\u20135 maturity scale and rolled up into a per-section average plus an overall average in the scorecard." + }, + "without_skill": { + "eval_id": "dataops-five-dimensions", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'Architecture' or contains 'architecture'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'architecture': Substring not found: 'architecture'; contains 'architecture': Substring not found: 'architecture'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Security' or contains 'security'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'security': Substring not found: 'security'; contains 'security': Substring not found: 'security'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Observability' or contains 'observability'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'observability': Substring not found: 'observability'; contains 'observability': Substring not found: 'observability'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Automation' or contains 'automation'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'automation': Substring not found: 'automation'; contains 'automation': Substring not found: 'automation'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Cost' or contains 'cost'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'cost': Substring not found: 'cost'; contains 'cost': Substring not found: 'cost'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.0, + "summary": "Without skill: 0% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 560, + "cache_read_input_tokens": 11297, + "cache_creation_input_tokens": 18548 + } + }, + "timing": { + "elapsed_seconds": 14.481381958001293 + }, + "raw_output": "I don't see a skill for a \"DataOps maturity assessment\" among the ones available in this session. Scanning the full list, the closest data/ops-adjacent skills are things like `builder-insights` (org/pipeline metrics), the pipeline-assistant family (pipeline health/friction), and `golden-path-recommendations` \u2014 but none of them defines a five-dimension DataOps maturity model.\n\nRather than guess at five dimensions and risk inventing them, I'd rather flag that there's no matching skill here. A couple of possibilities:\n\n- The skill may live in a plugin or directory that isn't loaded in this session.\n- It might go by a different name than \"DataOps maturity assessment.\"\n\nIf you can point me to the skill's name or where it's defined, I'll read it and list the exact five dimensions it scores. Want me to search the workspace/internal sources for a DataOps maturity skill?" + }, + "delta_pass_rate": 1.0 + }, + { + "eval_id": "dataops-scoring-scale", + "run_index": 0, + "with_skill": { + "eval_id": "dataops-scoring-scale", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '1-5' or contains '1 to 5' or contains '1 (lowest)'", + "passed": true, + "evidence": "OR satisfied by: contains '1-5' \u2014 Substring found: '1-5'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '26'", + "passed": true, + "evidence": "Substring found: '26'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 296, + "cache_read_input_tokens": 11306, + "cache_creation_input_tokens": 29168 + } + }, + "timing": { + "elapsed_seconds": 11.198131792014465 + }, + "raw_output": "The skill scores each question on a **1-5 maturity scale** (1 = lowest maturity, 5 = highest), and there are **26 questions** in total.\n\nBoth facts come straight from the skill definition \u2014 no AWS access needed. The scoring model states each question is scored \"1 (lowest) to 5 (highest),\" and the mandatory coverage rule enumerates all 26: Q3, Q4, Q5, Q6, Q7, Q8, Q11, Q12, Q13, Q14, Q19, Q20, Q23, Q24, Q25, Q26, Q29, Q30, Q31, Q32, Q33, Q36, Q37, Q38, Q39, Q40.\n\n(Note that some questions carry an auto-score ceiling of 3 \u2014 meaning control-plane signals alone can't push them above 3 without conversational confirmation \u2014 but the scale itself is still 1-5 for every question.)" + }, + "without_skill": { + "eval_id": "dataops-scoring-scale", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '1-5' or contains '1 to 5' or contains '1 (lowest)'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains '1-5': Substring not found: '1-5'; contains '1 to 5': Substring not found: '1 to 5'; contains '1 (lowest)': Substring not found: '1 (lowest)'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '26'", + "passed": true, + "evidence": "Substring found: '26'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.5, + "summary": "Without skill: 50% assertions passed", + "execution_metrics": { + "tool_calls": 4, + "token_counts": { + "input_tokens": 10, + "output_tokens": 1716, + "cache_read_input_tokens": 138032, + "cache_creation_input_tokens": 34466 + } + }, + "timing": { + "elapsed_seconds": 36.17559229201288 + }, + "raw_output": "I'll look at the available skills to find the one that scores questions on a maturity scale. Let me first check the working directory and locate the skill.The working directory is empty. The skill must be among the installed ones. Let me find which skill uses a maturity scale to score questions.Those are conversation logs. Let me find the actual skill definitions. There's a reference to a \"sample-agent-skill-eval\" path.There's a skill at `analytics-dataops-expertise`. Let me read its SKILL.md.I found it \u2014 the `analytics-dataops-expertise` skill (DataOps Maturity Assessment).\n\n## Answer\n\n**Maturity scale:** Each question is scored on a **1\u20135 maturity scale** (1 = lowest maturity, 5 = highest). For each question you pick the highest rating whose conditions are fully met by the observed signals; if signals fall between two ratings, you choose the lower one. Some questions carry an **auto-score ceiling of 3** \u2014 they can't be fully confirmed from control-plane API signals alone, so they cap at 3 unless conversational confirmation is available.\n\nSection and overall scores roll up as averages, with a status heuristic: **avg < 2 \u2192 \ud83d\udd34, 2\u20133.5 \u2192 \ud83d\udfe1, > 3.5 \u2192 \ud83d\udfe2**.\n\n**Total questions: 26.** They are:\n- **Architecture:** Q3, Q4, Q5, Q6, Q7, Q8\n- **Security & Governance:** Q29, Q30, Q31, Q32, Q33\n- **Incident Mgmt & Observability:** Q11, Q12, Q13, Q14, Q19, Q20\n- **Automation & Testing:** Q23, Q24, Q25, Q26\n- **Cost:** Q36, Q37, Q38, Q39, Q40\n\n(The numbering is non-contiguous \u2014 it skips up to Q40 \u2014 but the mandatory coverage rule requires exactly these 26 rows in the score matrix.)\n\nNo AWS access was needed for this; I answered by reading the skill's `SKILL.md`." + }, + "delta_pass_rate": 0.5 + }, + { + "eval_id": "dataops-auto-score-cap", + "run_index": 0, + "with_skill": { + "eval_id": "dataops-auto-score-cap", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '3'", + "passed": true, + "evidence": "Substring found: '3'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /cap|ceiling|maximum|max/", + "passed": true, + "evidence": "Regex /cap|ceiling|maximum|max/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /confirm|conversation/", + "passed": true, + "evidence": "Regex /confirm|conversation/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 923, + "cache_read_input_tokens": 11306, + "cache_creation_input_tokens": 29170 + } + }, + "timing": { + "elapsed_seconds": 18.828073709009914 + }, + "raw_output": "The cap is **3** (out of the 1\u20135 maturity scale).\n\n**Why it applies:** These questions can't be fully confirmed from control-plane API signals alone \u2014 *the existence of a resource is not the same as mature use of it.* The read-only APIs can tell you a thing exists (a dashboard, a Schema Registry, a Glue Workflow, an orchestration tool, a DQ ruleset, backup plans, etc.), but they can't confirm the operational maturity that a 4 or 5 requires \u2014 things like predictive/ML-driven behavior, enforced-on-producers policy, tested RPO/RTO with DR drills, automated remediation, formal SLA programs, or self-healing pipelines. Those higher ratings need **conversational confirmation** with the customer; they can't be inferred from a control-plane signal.\n\nSo the rule is: for a capped question, never auto-score above 3 from API signals, and state explicitly that ratings above the cap require conversational confirmation.\n\n**Which questions are capped at 3:** Q6, Q8, Q11, Q13, Q14, Q19, Q20, Q23, Q24, Q26, Q30, Q37, Q38, Q39.\n\nExamples of *why* each is capped:\n- **Q6 Capacity Planning** \u2014 auto-scaling policies are visible, but predictive/ML-based scaling + forecasting isn't.\n- **Q8 Disaster Recovery** \u2014 backups/PITR/cross-region copies are visible, but tested RPO/RTO and DR drills aren't.\n- **Q11 Workload Analytics** \u2014 a dashboard existing \u2260 active, KPI-driven operations.\n- **Q14 Drift Detection** \u2014 a Schema Registry/DQ ruleset existing \u2260 enforced on producers.\n- **Q30 Lineage** \u2014 Glue Workflows imply *some* lineage, but automated end-to-end lineage tooling is a data-plane/tooling property.\n- **Q37 Chargeback/Showback** \u2014 cost-allocation tags are visible, but an automated chargeback program isn't.\n\nBy contrast, the **no-cap** questions (Q3, Q4, Q5, Q7, Q12, Q25, Q29, Q31, Q32, Q33, Q36, Q40) can be scored the full 1\u20135 from signals, because the API evidence directly reflects the maturity being measured (e.g., encryption coverage, Multi-AZ topology, tagging percentage, presence of stre" + }, + "without_skill": { + "eval_id": "dataops-auto-score-cap", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '3'", + "passed": false, + "evidence": "Substring not found: '3'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /cap|ceiling|maximum|max/", + "passed": false, + "evidence": "Regex /cap|ceiling|maximum|max/ did not match", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /confirm|conversation/", + "passed": false, + "evidence": "Regex /confirm|conversation/ did not match", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.0, + "summary": "Without skill: 0% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 180.00441537500592 + }, + "raw_output": "" + }, + "delta_pass_rate": 1.0 + }, + { + "eval_id": "dataops-required-input", + "run_index": 0, + "with_skill": { + "eval_id": "dataops-required-input", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'account'", + "passed": true, + "evidence": "Substring found: 'account'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'region'", + "passed": true, + "evidence": "Substring found: 'region'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /ask|missing|do not assume|prompt/", + "passed": true, + "evidence": "Regex /ask|missing|do not assume|prompt/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 343, + "cache_read_input_tokens": 11306, + "cache_creation_input_tokens": 29160 + } + }, + "timing": { + "elapsed_seconds": 13.656385667010909 + }, + "raw_output": "The skill requires an **account ID and a region** (e.g., `123456789012` / `eu-west-1`) before it can run. The scorer is account + region scoped, and the entire assessment runs against exactly that pair.\n\n**If the input is missing:** the skill says to **ask the user for it \u2014 do not assume.** It won't guess a default account or region and won't proceed on a partial input.\n\nRelevant excerpts:\n\n- Under *Execution Flow \u2192 Input*: \"**Required:** account ID + region... If the account or region is missing, ask for it \u2014 do not assume. The assessment runs against exactly the account + region provided.\"\n- Step 1 of the flow is literally \"Confirm target account ID + region with the user.\"\n\nSo it's a hard gate: no account+region, no run \u2014 it stops and prompts you rather than making assumptions or picking a default." + }, + "without_skill": { + "eval_id": "dataops-required-input", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'account'", + "passed": false, + "evidence": "Substring not found: 'account'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'region'", + "passed": false, + "evidence": "Substring not found: 'region'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /ask|missing|do not assume|prompt/", + "passed": true, + "evidence": "Regex /ask|missing|do not assume|prompt/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.3333333333333333, + "summary": "Without skill: 33% assertions passed", + "execution_metrics": { + "tool_calls": 20, + "token_counts": { + "input_tokens": 42, + "output_tokens": 8554, + "cache_read_input_tokens": 911762, + "cache_creation_input_tokens": 47905 + } + }, + "timing": { + "elapsed_seconds": 162.5909753330052 + }, + "raw_output": "I'll look at the skill in the working directory to answer this.The working directory is empty, so \"the skill\" must be one of the available skills. Let me find the skill definitions on disk to see which one this refers to.The referenced paths don't exist. Let me locate the actual plugin skill files, since the question targets a specific skill.This is a paired skill-eval (a \"without\" variant \u2014 the skill is deliberately not loaded into my working directory). Let me check the eval harness to see which skill this question is about.This eval is about the `skill-eval` skill in that repo. Let me read its SKILL.md to answer precisely.Every command takes a skill path as input. Let me check the CLI reference and implementation to see exactly what happens when that input is missing.The required input is a skill directory path. Let me confirm the missing-input behavior in the CLI implementation.I found the skill this eval targets \u2014 it's the **`skill-eval`** skill (an agent-skill evaluation framework). Here's what it requires and how it handles missing input.\n\n## Required input\n\nEvery `skill-eval` command takes a **path to a skill directory** as its input \u2014 and that directory must contain a `SKILL.md` file at its root (per the framework, \"the only required file\"). Concretely, at the CLI level (`skill_eval/cli.py`):\n\n- `audit` \u2014 one or more `skill_path` positionals (`nargs=\"+\"`)\n- `functional`, `trigger`, `report`, `init`, `snapshot`, `regression`, `lifecycle` \u2014 a single required `skill_path`\n- `compare` \u2014 two required paths (`skill_a`, `skill_b`)\n\nSome commands need **additional inputs** beyond the path:\n- `functional` needs `evals/evals.json`\n- `trigger` needs `evals/eval_queries.json`\n- `compare` needs an evals file for both skills\n\n## What it does when input is missing\n\nThe behavior differs by *which* input is missing:\n\n**No subcommand at all** \u2192 prints full help text and exits `0` (`cli.py:267`).\n\n**Subcommand given but `skill_path` omitted** \u2192 argparse rejects it with a usag" + }, + "delta_pass_rate": 0.6666666666666667 + }, + { + "eval_id": "dataops-control-plane-only", + "run_index": 0, + "with_skill": { + "eval_id": "dataops-control-plane-only", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'control-plane' or contains 'control plane'", + "passed": true, + "evidence": "OR satisfied by: contains 'control-plane' \u2014 Substring found: 'control-plane'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /read-only|read only/", + "passed": true, + "evidence": "Regex /read-only|read only/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /internal|data-plane|data plane/", + "passed": true, + "evidence": "Regex /internal|data-plane|data plane/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 878, + "cache_read_input_tokens": 11306, + "cache_creation_input_tokens": 29162 + } + }, + "timing": { + "elapsed_seconds": 18.66560979100177 + }, + "raw_output": "This is answerable straight from the skill definition \u2014 no AWS calls needed.\n\n## Data sources the skill **uses**\n\n**Only read-only AWS control-plane APIs**, scoped to a single account + region. Absence of a service is itself a valid (low-maturity) signal. The APIs span these services:\n\n- **Architecture:** `glue` (databases, jobs, crawlers, workflows), `lakeformation`, `kinesis`, `firehose`, `kinesisanalyticsv2` (Flink), `kafka`/MSK, `dms`, `dynamodb` (+ Streams), `application-autoscaling`, `autoscaling`, `rds`, `es`/OpenSearch, `redshift`, `backup`\n- **Security & Governance:** `macie2`, `kms`, `s3` (config/encryption/lifecycle), `iam`, `config`, `lakeformation` permissions/cell filters\n- **Observability:** `cloudwatch` (alarms, dashboards, metric data), `sns`, `events`/EventBridge, `xray`, `rum`, `quicksight`, `glue` (Schema Registry, Data Quality)\n- **Automation & Testing:** `cloudformation`, `codepipeline`, `codecommit`, `mwaa`, `states`/Step Functions, `ssm`, `lambda`\n- **Cost:** `resourcegroupstaggingapi`, `ce`/Cost Explorer (in `us-east-1`), `cost-optimization-hub`, `emr`\n\n## What it **explicitly avoids**\n\n- **No data-plane access** \u2014 no queries against catalogs, tables, or indices; only control-plane describe/list/get calls.\n- **No AWS-internal data sources** \u2014 it does **not** use support-case history, internal customer records, or any AWS-internal data. It won't infer or fabricate such signals.\n- **No writes / no destructive operations** \u2014 read-only throughout (list/describe/get).\n- **No external packages, no IDE workspace** \u2014 CLI-agent compatible, account-scoped only.\n- **No fabrication** \u2014 if a signal can't be obtained from account control-plane APIs, the skill says so rather than inventing it.\n\nA related design consequence: because control-plane existence \u2260 mature usage, several questions are **capped at score 3** from API signals alone \u2014 anything higher requires conversational confirmation rather than a guessed data source." + }, + "without_skill": { + "eval_id": "dataops-control-plane-only", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'control-plane' or contains 'control plane'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'control-plane': Substring not found: 'control-plane'; contains 'control plane': Substring not found: 'control plane'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /read-only|read only/", + "passed": false, + "evidence": "Regex /read-only|read only/ did not match", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /internal|data-plane|data plane/", + "passed": false, + "evidence": "Regex /internal|data-plane|data plane/ did not match", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.0, + "summary": "Without skill: 0% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 180.00792962498963 + }, + "raw_output": "" + }, + "delta_pass_rate": 1.0 + }, + { + "eval_id": "dataops-missing-service-signal", + "run_index": 0, + "with_skill": { + "eval_id": "dataops-missing-service-signal", + "run_index": 0, + "assertion_results": [ + { + "text": "matches regex /score-1|score 1|not present|score of 1/", + "passed": true, + "evidence": "Regex /score-1|score 1|not present|score of 1/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'SKIPPED' or contains 'skipped' or contains 'skip'", + "passed": true, + "evidence": "OR satisfied by: contains 'skipped' \u2014 Substring found: 'skipped'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /denied|AccessDenied|IAM/", + "passed": true, + "evidence": "Regex /denied|AccessDenied|IAM/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 496, + "cache_read_input_tokens": 11306, + "cache_creation_input_tokens": 29174 + } + }, + "timing": { + "elapsed_seconds": 14.55529945800663 + }, + "raw_output": "The skill treats all three of those as the **same \"resource not present\" condition \u2014 a score-1 signal, never a blocker and never SKIPPED**. Specifically:\n\n- **API errors / unavailable in region / empty results** \u2192 \"treat as 'not present' (a score-1 signal), never as a blocker.\" A region with no Kinesis/MSK/Flink genuinely means \"no real-time processing,\" so Q4 = 1, and so on.\n\n- **The one exception is IAM denial.** If an API returns `AccessDeniedException`:\n - If the question has *other* signals available, score from those and note the gap.\n - Only mark a question **SKIPPED** if *every* API that question needs is denied by IAM \u2014 and then it must state which permission is missing.\n\n- **Cost Explorer special case:** `ce.*` must be enabled and is called in `us-east-1`. If unavailable/AccessDenied, Q37/Q39 are scored from tagging or other signals with the CE gap noted \u2014 not skipped (unless there's no signal at all).\n\nThe guiding principle: absence of a service is itself a valid low-maturity signal, so a missing or erroring resource pushes the score toward 1 rather than removing the question from the assessment. And it never fabricates a signal it can't obtain from control-plane APIs." + }, + "without_skill": { + "eval_id": "dataops-missing-service-signal", + "run_index": 0, + "assertion_results": [ + { + "text": "matches regex /score-1|score 1|not present|score of 1/", + "passed": true, + "evidence": "Regex /score-1|score 1|not present|score of 1/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'SKIPPED' or contains 'skipped' or contains 'skip'", + "passed": true, + "evidence": "OR satisfied by: contains 'skipped' \u2014 Substring found: 'skipped'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /denied|AccessDenied|IAM/", + "passed": true, + "evidence": "Regex /denied|AccessDenied|IAM/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "Without skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 9, + "token_counts": { + "input_tokens": 20, + "output_tokens": 4995, + "cache_read_input_tokens": 332032, + "cache_creation_input_tokens": 41385 + } + }, + "timing": { + "elapsed_seconds": 101.93233895802405 + }, + "raw_output": "I'll look for the skill in question \u2014 the phrasing suggests there's a specific skill that enumerates AWS resources across regions. Let me find it.The working directory is empty. Let me find the skill that deals with AWS APIs across regions.There are multiple AWS-related skills. Let me see the full list to identify which one enumerates resources across regions.Let me search across the skills for the error-handling / region / no-resources language.The question likely comes from an eval benchmark. Let me search the eval JSON files for this exact question to find the target skill.The target is the **analytics-dataops-expertise** skill. Let me read it and its references to answer accurately.The skill is **`analytics-dataops-expertise`** (a read-only DataOps maturity scorer). Here's exactly how it treats each of those three conditions \u2014 the rule is stated consistently in its Execution Flow (step 2), Scoring model, and Error Handling section.\n\n## The core principle\nAbsence of a signal is *itself* a valid finding \u2014 a **score-1 (lowest-maturity) signal**, never a blocker. The skill is scoring maturity 1\u20135, so \"the customer has no real-time processing here\" is a legitimate answer, not a failure to assess. It never fabricates a signal it couldn't obtain.\n\n## The three cases\n\n**1. API errors** \u2014 it depends on the error type:\n- **`AccessDeniedException` (IAM):** If the question has *other* usable signals, score from those and note the permission gap. A question is only marked **SKIPPED** if **every** API it needs is denied \u2014 and even then it must state which permission is missing. So IAM denial is the *only* thing that produces a SKIP, never a score.\n- **Any other error:** treated as \"resource not present\" \u2192 a score-1 signal, \"never as a blocker\" (SKILL.md:76\u201377, 96\u201398).\n\n**2. Unavailable in the region:** treated identically to empty \u2014 \"resource not present,\" a score-1 signal, *not* an error (SKILL.md:420\u2013421). Example given: a region with no Kinesis/MSK/Flink genuinely means \"n" + }, + "delta_pass_rate": 0.0 + }, + { + "eval_id": "dataops-cost-explorer-prereq", + "run_index": 0, + "with_skill": { + "eval_id": "dataops-cost-explorer-prereq", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'Cost Explorer' or contains 'cost explorer'", + "passed": true, + "evidence": "OR satisfied by: contains 'cost explorer' \u2014 Substring found: 'cost explorer'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'us-east-1'", + "passed": true, + "evidence": "Substring found: 'us-east-1'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /enabled|enable/", + "passed": true, + "evidence": "Regex /enabled|enable/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 478, + "cache_read_input_tokens": 11306, + "cache_creation_input_tokens": 29170 + } + }, + "timing": { + "elapsed_seconds": 13.914667500008363 + }, + "raw_output": "The two Cost Explorer\u2013dependent questions are **Q37 (Chargeback / Showback)** and **Q39 (Data Processing Cost)**.\n\nRequirements for both:\n\n- **Cost Explorer must be enabled** in the account. The relevant APIs are `ce.GetCostAndUsage` (and `ce.GetTags` for Q37), plus `cost-explorer` cost-allocation tag data.\n- **It must be called in the `us-east-1` region** \u2014 Cost Explorer is a global service whose endpoint lives in `us-east-1`.\n- If Cost Explorer is **not enabled or returns AccessDenied**, you do **not** skip these questions. Instead, score them from other available signals (e.g., tagging signals from `resourcegroupstaggingapi`, plus `emr`/`glue` sizing for Q39) and **note the CE gap** in the report. A question is only marked SKIPPED if *every* API it needs is denied by IAM.\n\nBoth are also **capped at a maximum auto-score of 3** \u2014 ratings of 4\u20135 (automated chargeback for Q37; ML-based/zero-waste optimization for Q39) require conversational confirmation, since control-plane signals alone can't confirm them." + }, + "without_skill": { + "eval_id": "dataops-cost-explorer-prereq", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'Cost Explorer' or contains 'cost explorer'", + "passed": true, + "evidence": "OR satisfied by: contains 'cost explorer' \u2014 Substring found: 'cost explorer'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'us-east-1'", + "passed": true, + "evidence": "Substring found: 'us-east-1'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /enabled|enable/", + "passed": true, + "evidence": "Regex /enabled|enable/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "Without skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 4, + "token_counts": { + "input_tokens": 6, + "output_tokens": 2159, + "cache_read_input_tokens": 71664, + "cache_creation_input_tokens": 23356 + } + }, + "timing": { + "elapsed_seconds": 40.1935078749957 + }, + "raw_output": "I'll search for the reference material about these cost questions.I checked the working directory and the Claude config/skill directories, and the only files that mention \"Cost Explorer\" are session transcript logs (`.jsonl`) \u2014 echoes of this same question from other eval runs, not actual reference material. There is no document, skill, or SOP available to me here that defines \"the two cost questions that rely on Cost Explorer.\"\n\nSo I can't tell you *which* two questions those are without inventing them, and I won't guess at that.\n\nWhat I can state factually:\n\n- **Region:** The AWS Cost Explorer API (`ce`) is a global service whose endpoint lives in **`us-east-1`**. Any code calling `GetCostAndUsage` / `GetCostForecast` etc. must target `us-east-1` regardless of where the workload runs.\n- **Typical requirements** for Cost Explorer calls (in general): Cost Explorer must be enabled on the account/payer, and the caller needs `ce:Get*` IAM permissions (e.g. `ce:GetCostAndUsage`, `ce:GetCostForecast`). Data is delayed (~24h) and there's a per-request charge for paginated API calls.\n\nIf \"the two cost questions\" comes from a specific skill or runbook you're working from, point me at that file (or its name) and I'll answer precisely from it rather than from general knowledge." + }, + "delta_pass_rate": 0.0 + }, + { + "eval_id": "dataops-status-heuristic", + "run_index": 0, + "with_skill": { + "eval_id": "dataops-status-heuristic", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '2'", + "passed": true, + "evidence": "Substring found: '2'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '3.5'", + "passed": true, + "evidence": "Substring found: '3.5'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 112, + "cache_read_input_tokens": 11306, + "cache_creation_input_tokens": 29167 + } + }, + "timing": { + "elapsed_seconds": 8.958538457984105 + }, + "raw_output": "The status thresholds are defined in the Output Format's status heuristic:\n\n- \ud83d\udd34 **Red**: avg < 2\n- \ud83d\udfe1 **Yellow**: avg 2\u20133.5\n- \ud83d\udfe2 **Green**: avg > 3.5\n\nThese apply to each section's average score (mean of its question scores, 1\u20135) in the summary table." + }, + "without_skill": { + "eval_id": "dataops-status-heuristic", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '2'", + "passed": false, + "evidence": "Substring not found: '2'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '3.5'", + "passed": false, + "evidence": "Substring not found: '3.5'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.0, + "summary": "Without skill: 0% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 180.00375041700318 + }, + "raw_output": "" + }, + "delta_pass_rate": 1.0 + } + ], + "run_summary": { + "with_skill": { + "mean_pass_rate": 1.0, + "stddev_pass_rate": 0.0, + "mean_tokens": 463.9, + "mean_input_tokens": 2.0, + "mean_output_tokens": 463.9, + "mean_total_tokens": 465.9, + "mean_tool_calls": 0.0 + }, + "without_skill": { + "mean_pass_rate": 0.3542, + "stddev_pass_rate": 0.4404, + "mean_tokens": 2248.0, + "mean_input_tokens": 10.0, + "mean_output_tokens": 2248.0, + "mean_total_tokens": 2258.0, + "mean_tool_calls": 4.6 + }, + "delta": { + "pass_rate": 0.6458, + "tokens": -1784.1, + "total_tokens": -1792.1, + "input_tokens": -8.0, + "tool_calls": -4.6 + }, + "cost_efficiency": { + "quality_delta": 0.6458, + "cost_delta_pct": -79.4, + "classification": "PARETO_BETTER", + "emoji": "\ud83d\udfe2", + "description": "Skill improves quality while reducing cost" + }, + "estimated_cost": { + "with_skill_per_run": { + "input_cost": 6e-06, + "output_cost": 0.006958, + "total_cost": 0.006964, + "model": "sonnet", + "currency": "USD" + }, + "without_skill_per_run": { + "input_cost": 3e-05, + "output_cost": 0.03372, + "total_cost": 0.03375, + "model": "sonnet", + "currency": "USD" + }, + "per_eval_pair": 0.040714, + "total_runs": 8, + "total_cost": 0.3257, + "model": "sonnet", + "currency": "USD" + } + }, + "scores": { + "outcome": 1.0, + "process": 0.0, + "style": 1.0, + "efficiency": 1.0, + "overall": 0.75 + }, + "passed": true +} \ No newline at end of file diff --git a/skills/analytics-dataops-expertise/evals/eval_queries.json b/skills/analytics-dataops-expertise/evals/eval_queries.json new file mode 100644 index 0000000..321ddbe --- /dev/null +++ b/skills/analytics-dataops-expertise/evals/eval_queries.json @@ -0,0 +1,9 @@ +[ + {"query": "Which skill would help me assess the DataOps maturity of my AWS data platform? Just name it; do not run it.", "should_trigger": true}, + {"query": "Is there a skill for reviewing our data architecture, governance posture, pipeline orchestration, and data cost optimization? Answer with the skill name; do not execute it.", "should_trigger": true}, + {"query": "Name the skill that scores analytics maturity across Glue, Kinesis, DMS, and observability on a 1-5 scale. Do not run any assessment.", "should_trigger": true}, + {"query": "Write a Python script that sorts a list of numbers", "should_trigger": false}, + {"query": "What's the weather forecast for Sydney this weekend?", "should_trigger": false}, + {"query": "Create a CloudFormation template for an S3 bucket", "should_trigger": false}, + {"query": "Review the health and configuration of my Amazon OpenSearch domain", "should_trigger": false} +] diff --git a/skills/analytics-dataops-expertise/evals/evals.json b/skills/analytics-dataops-expertise/evals/evals.json new file mode 100644 index 0000000..bf6b02d --- /dev/null +++ b/skills/analytics-dataops-expertise/evals/evals.json @@ -0,0 +1,90 @@ +[ + { + "id": "dataops-five-dimensions", + "prompt": "According to the skill, list the five dimensions it scores for a DataOps maturity assessment. No AWS access required.", + "expected_output": "Lists Architecture; Security & Governance; Incident Management & Observability; Automation & Testing; and Cost.", + "files": [], + "assertions": [ + "contains 'Architecture' or contains 'architecture'", + "contains 'Security' or contains 'security'", + "contains 'Observability' or contains 'observability'", + "contains 'Automation' or contains 'automation'", + "contains 'Cost' or contains 'cost'" + ] + }, + { + "id": "dataops-scoring-scale", + "prompt": "What maturity scale does the skill use to score each question, and how many questions are there in total? No AWS access required.", + "expected_output": "Each question is scored 1-5; there are 26 questions total.", + "files": [], + "assertions": [ + "contains '1-5' or contains '1 to 5' or contains '1 (lowest)'", + "contains '26'" + ] + }, + { + "id": "dataops-auto-score-cap", + "prompt": "The skill caps some questions at a maximum auto-score. What is that cap and why does it apply? No AWS access required.", + "expected_output": "Certain questions are capped at 3 because existence of a resource does not confirm mature use; a higher score requires conversational confirmation.", + "files": [], + "assertions": [ + "contains '3'", + "matches regex /cap|ceiling|maximum|max/", + "matches regex /confirm|conversation/" + ] + }, + { + "id": "dataops-required-input", + "prompt": "What input does the skill require before it can run, and what does it do if that input is missing? No AWS access required.", + "expected_output": "Requires an account ID and a region; if either is missing it asks the user rather than assuming.", + "files": [], + "assertions": [ + "contains 'account'", + "contains 'region'", + "matches regex /ask|missing|do not assume|prompt/" + ] + }, + { + "id": "dataops-control-plane-only", + "prompt": "What data sources does the skill use, and what does it explicitly avoid? No AWS access required.", + "expected_output": "Uses only read-only account control-plane APIs; no data-plane access and no AWS-internal data sources such as support cases.", + "files": [], + "assertions": [ + "contains 'control-plane' or contains 'control plane'", + "matches regex /read-only|read only/", + "matches regex /internal|data-plane|data plane/" + ] + }, + { + "id": "dataops-missing-service-signal", + "prompt": "How does the skill treat an AWS API that errors, is unavailable in the region, or returns no resources? No AWS access required.", + "expected_output": "Treats it as 'resource not present', which is a score-1 signal, not a blocker; a question is only SKIPPED when every API it needs is denied by IAM.", + "files": [], + "assertions": [ + "matches regex /score-1|score 1|not present|score of 1/", + "contains 'SKIPPED' or contains 'skipped' or contains 'skip'", + "matches regex /denied|AccessDenied|IAM/" + ] + }, + { + "id": "dataops-cost-explorer-prereq", + "prompt": "What is required for the two cost questions that rely on Cost Explorer, and in which region is it called? No AWS access required.", + "expected_output": "Cost Explorer must be enabled and is called in us-east-1; if unavailable, those questions score from tagging signals with the gap noted.", + "files": [], + "assertions": [ + "contains 'Cost Explorer' or contains 'cost explorer'", + "contains 'us-east-1'", + "matches regex /enabled|enable/" + ] + }, + { + "id": "dataops-status-heuristic", + "prompt": "What status thresholds does the skill's summary table use to assign red / yellow / green to each section? No AWS access required.", + "expected_output": "Average below 2 is red, 2 to 3.5 is yellow, above 3.5 is green.", + "files": [], + "assertions": [ + "contains '2'", + "contains '3.5'" + ] + } +] diff --git a/skills/analytics-dataops-expertise/evals/functional_report.json b/skills/analytics-dataops-expertise/evals/functional_report.json new file mode 100644 index 0000000..b3bd417 --- /dev/null +++ b/skills/analytics-dataops-expertise/evals/functional_report.json @@ -0,0 +1,834 @@ +{ + "skill_name": "analytics-dataops-expertise", + "skill_path": "/Users/prasadnu/Dante/OpenSearch-OpsReview-KiroPower/external-repos/tools-for-devops-agent/skills/analytics-dataops-expertise", + "eval_count": 8, + "runs_per_eval": 1, + "metadata": { + "timestamp": "2026-08-28T11:15:40Z" + }, + "runs": [ + { + "eval_id": "dataops-five-dimensions", + "run_index": 0, + "with_skill": { + "eval_id": "dataops-five-dimensions", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'Architecture' or contains 'architecture'", + "passed": true, + "evidence": "OR satisfied by: contains 'architecture' \u2014 Substring found: 'architecture'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Security' or contains 'security'", + "passed": true, + "evidence": "OR satisfied by: contains 'security' \u2014 Substring found: 'security'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Observability' or contains 'observability'", + "passed": true, + "evidence": "OR satisfied by: contains 'observability' \u2014 Substring found: 'observability'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Automation' or contains 'automation'", + "passed": true, + "evidence": "OR satisfied by: contains 'automation' \u2014 Substring found: 'automation'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Cost' or contains 'cost'", + "passed": true, + "evidence": "OR satisfied by: contains 'cost' \u2014 Substring found: 'cost'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 112, + "cache_read_input_tokens": 11306, + "cache_creation_input_tokens": 29169 + } + }, + "timing": { + "elapsed_seconds": 9.51375654200092 + }, + "raw_output": "The five dimensions the skill scores are:\n\n1. **Architecture**\n2. **Security & Governance**\n3. **Incident Management & Observability**\n4. **Automation & Testing**\n5. **Cost**\n\nEach is scored on a 1\u20135 maturity scale, then rolled up to a per-section average and an overall average across all 26 questions." + }, + "without_skill": { + "eval_id": "dataops-five-dimensions", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'Architecture' or contains 'architecture'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'architecture': Substring not found: 'architecture'; contains 'architecture': Substring not found: 'architecture'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Security' or contains 'security'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'security': Substring not found: 'security'; contains 'security': Substring not found: 'security'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Observability' or contains 'observability'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'observability': Substring not found: 'observability'; contains 'observability': Substring not found: 'observability'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Automation' or contains 'automation'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'automation': Substring not found: 'automation'; contains 'automation': Substring not found: 'automation'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Cost' or contains 'cost'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'cost': Substring not found: 'cost'; contains 'cost': Substring not found: 'cost'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.0, + "summary": "Without skill: 0% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 618, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 29843 + } + }, + "timing": { + "elapsed_seconds": 16.788515374995768 + }, + "raw_output": "I don't have a skill matching that description. I looked through all the skills available in this session, and none of them is a \"DataOps maturity assessment\" skill or scores maturity across five dimensions. The available skills cover things like Amazon build/pipeline tooling (Brazil, CRUX, Apollo/pipeline troubleshooting, gatekeeper dependency analysis), data visualization, browser/game testing, config/keybindings, and the Claude API reference \u2014 none of them define a DataOps maturity rubric.\n\nSo I can't list \"the five dimensions,\" because there's no such skill here to read them from, and I won't invent them (a fabricated maturity rubric would be worse than useless).\n\nA couple of ways forward:\n- If the skill lives elsewhere (a repo file, a wiki page, a package), point me at it and I'll read it and list the dimensions verbatim.\n- If you're thinking of a specific named skill, tell me the name and I'll check whether it's actually available.\n- If you just want a well-grounded DataOps maturity model, I can sketch one from common industry frameworks \u2014 but I'll clearly label it as general guidance, not \"the skill's\" dimensions.\n\nWhich would you like?" + }, + "delta_pass_rate": 1.0 + }, + { + "eval_id": "dataops-scoring-scale", + "run_index": 0, + "with_skill": { + "eval_id": "dataops-scoring-scale", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '1-5' or contains '1 to 5' or contains '1 (lowest)'", + "passed": true, + "evidence": "OR satisfied by: contains '1-5' \u2014 Substring found: '1-5'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '26'", + "passed": true, + "evidence": "Substring found: '26'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 269, + "cache_read_input_tokens": 11306, + "cache_creation_input_tokens": 29164 + } + }, + "timing": { + "elapsed_seconds": 13.227103167009773 + }, + "raw_output": "The skill scores each question on a **1-to-5 maturity scale** (1 = lowest maturity, 5 = highest), picking the highest rating whose conditions are fully met by the observed signals \u2014 and choosing the lower of two when signals fall in between.\n\nThere are **26 questions** in total: Q3, Q4, Q5, Q6, Q7, Q8, Q11, Q12, Q13, Q14, Q19, Q20, Q23, Q24, Q25, Q26, Q29, Q30, Q31, Q32, Q33, Q36, Q37, Q38, Q39, and Q40.\n\nOne nuance worth noting: some questions have an **auto-score ceiling of 3** (they can't be fully confirmed from control-plane API signals alone), while others have no cap and can score the full 1-5 from signals." + }, + "without_skill": { + "eval_id": "dataops-scoring-scale", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '1-5' or contains '1 to 5' or contains '1 (lowest)'", + "passed": true, + "evidence": "OR satisfied by: contains '1-5' \u2014 Substring found: '1-5'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '26'", + "passed": true, + "evidence": "Substring found: '26'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "Without skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 9, + "token_counts": { + "input_tokens": 20, + "output_tokens": 3850, + "cache_read_input_tokens": 413699, + "cache_creation_input_tokens": 58967 + } + }, + "timing": { + "elapsed_seconds": 173.24328687501838 + }, + "raw_output": "I'll look for the skill that uses a maturity scale to score questions. Let me find the skill files on disk.These are conversation logs. The question refers to a specific skill\u2014likely in the working directory or a related repo. Let me look.The working directory is empty. Let me search more broadly for the skill files that describe a maturity scale and scored questions.This is about a DataOps pre-assessment skill. Let me read its files.Let me read the dataops-preassessment skill's main file.The SKILL.md says 26 questions scored 1-5. Let me confirm what the 1-5 levels are named by checking the scoring engine.Let me check the scoring config and report for the named 1-5 levels.Let me see the maturity band labels in report.py around line 290-320.I have the answer. No AWS access was needed\u2014this is all in the skill's files.\n\n## Answer\n\n**Maturity scale:** The `dataops-preassessment` skill scores **each question on a deterministic 1\u20135 maturity scale** (1 = lowest maturity, 5 = highest), driven purely from live account API signals. Each question has a per-level rubric in `references/maturity_scoring/assessment/config.py` (e.g., level 5 = \"advanced lakehouse/mesh with federated governance\u2026 intelligent optimization\"). An API that errors or isn't onboarded is treated as a score-1 signal.\n\nThe per-question 1\u20135 scores roll up to an overall average that maps to four named **maturity bands** (in `assessment/report.py`):\n\n| Overall average | Band |\n|---|---|\n| \u2265 4 | Advanced |\n| 3 \u2013 <4 | Established |\n| 2 \u2013 <3 | Developing |\n| < 2 | Initial |\n\n**Total number of questions:** **26 questions.** This is the account-scoped maturity scorer (Step 10), the vendored engine at `references/maturity_scoring/`. It's optional\u2014it runs only when both an `account_id` and a `region` are supplied; otherwise only the customer-scoped ORA briefing is produced.\n\n(Note: the question collector files are numbered up to `q40`, but the actual scored set is 26 questions, as stated consistently in `SKILL.md`, the" + }, + "delta_pass_rate": 0.0 + }, + { + "eval_id": "dataops-auto-score-cap", + "run_index": 0, + "with_skill": { + "eval_id": "dataops-auto-score-cap", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '3'", + "passed": true, + "evidence": "Substring found: '3'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /cap|ceiling|maximum|max/", + "passed": true, + "evidence": "Regex /cap|ceiling|maximum|max/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /confirm|conversation/", + "passed": true, + "evidence": "Regex /confirm|conversation/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 571, + "cache_read_input_tokens": 11306, + "cache_creation_input_tokens": 29172 + } + }, + "timing": { + "elapsed_seconds": 15.123123833996942 + }, + "raw_output": "**The cap is 3** (a maximum auto-score of 3 out of 5).\n\nIt applies to questions where the control-plane APIs can confirm that a capability *exists*, but **cannot confirm it's used maturely** \u2014 as the skill puts it, \"existence of a resource \u2260 mature use of it.\" A read-only control-plane check can see that a dashboard, X-Ray group, orchestration tool, or DQ ruleset is *present*, but it can't verify the operational maturity that ratings 4\u20135 require (e.g., KPI-driven ops, enforced-on-producers schema validation, tested RPO/RTO, automated chargeback, self-healing pipelines). Those top ratings need **conversational confirmation** with the customer.\n\n**Capped-at-3 questions (14):** Q6, Q8, Q11, Q13, Q14, Q19, Q20, Q23, Q24, Q26, Q30, Q37, Q38, Q39.\n\nThe rule is applied *exactly*: for a capped question you never score above 3 from API signals alone, and you must state that a higher rating requires conversational confirmation. The remaining \"no cap\" questions (Q3, Q4, Q5, Q7, Q12, Q25, Q29, Q31, Q32, Q33, Q36, Q40) can be scored the full 1\u20135 from signals, because the API evidence there is sufficient to distinguish mature from immature use (e.g., encryption coverage, Multi-AZ topology, tagging percentage, version currency)." + }, + "without_skill": { + "eval_id": "dataops-auto-score-cap", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '3'", + "passed": false, + "evidence": "Substring not found: '3'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /cap|ceiling|maximum|max/", + "passed": false, + "evidence": "Regex /cap|ceiling|maximum|max/ did not match", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /confirm|conversation/", + "passed": false, + "evidence": "Regex /confirm|conversation/ did not match", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.0, + "summary": "Without skill: 0% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 180.00535158297862 + }, + "raw_output": "" + }, + "delta_pass_rate": 1.0 + }, + { + "eval_id": "dataops-required-input", + "run_index": 0, + "with_skill": { + "eval_id": "dataops-required-input", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'account'", + "passed": true, + "evidence": "Substring found: 'account'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'region'", + "passed": true, + "evidence": "Substring found: 'region'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /ask|missing|do not assume|prompt/", + "passed": true, + "evidence": "Regex /ask|missing|do not assume|prompt/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 2, + "output_tokens": 346, + "cache_read_input_tokens": 11306, + "cache_creation_input_tokens": 29164 + } + }, + "timing": { + "elapsed_seconds": 15.599186042003566 + }, + "raw_output": "The skill requires **an account ID and a region** (it's account + region scoped) \u2014 for example `123456789012` / `eu-west-1`.\n\n**If that input is missing:** the skill instructs you to **ask the user for it and not assume**. Per the Execution Flow:\n\n> \"If the account or region is missing, ask for it \u2014 do not assume. The assessment runs against exactly the account + region provided.\"\n\nAnd step 1 of the execution steps is explicitly: *\"Confirm target account ID + region with the user.\"*\n\nSo the required input is the target account ID plus region, and the defined behavior on a missing value is to stop and prompt for it rather than guessing or picking a default. (Everything else \u2014 the read-only AWS profile/permissions \u2014 is a prerequisite for actually running the checks, but the account ID + region is the specific input the skill will block on and ask for.)" + }, + "without_skill": { + "eval_id": "dataops-required-input", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'account'", + "passed": false, + "evidence": "Substring not found: 'account'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'region'", + "passed": false, + "evidence": "Substring not found: 'region'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /ask|missing|do not assume|prompt/", + "passed": true, + "evidence": "Regex /ask|missing|do not assume|prompt/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.3333333333333333, + "summary": "Without skill: 33% assertions passed", + "execution_metrics": { + "tool_calls": 6, + "token_counts": { + "input_tokens": 12, + "output_tokens": 3649, + "cache_read_input_tokens": 164601, + "cache_creation_input_tokens": 29570 + } + }, + "timing": { + "elapsed_seconds": 70.2946092079801 + }, + "raw_output": "I'll look at the skill in the working directory to answer this.The working directory is empty. The question refers to \"the skill\" but doesn't name one, and I don't have a specific skill file in this project. Let me check where the skills are defined to find the one you mean.There are two skills installed here (`browser-automation` and `game-development`), and your question isn't tied to a named one \u2014 but the phrasing \"requires input before it can run, and what it does if that input is missing\" maps cleanly onto **`game-development`**, so I'll answer that and note `browser-automation` too.\n\n## `game-development`\n\n**Required input:** a manifest file, `.codegpt-game.json`, in the project. It holds everything engine-specific \u2014 build command, install target, launch command, log paths, the `ready` pattern, RCON config, etc. Machine-specific values are supplied through `${VARS}` (resolved from the manifest's `vars` block first, then the environment). You bootstrap it with `node