Overview

Your agent's persona was never meant to carry your entire domain knowledge. As teams bring AI into real data engineering workflows—handling shifting ingestion feeds, inferring schemas, and standardizing lakehouse partitions—a single system instructions file quickly becomes a tangled mess of identity, procedure, and governance. It becomes difficult to maintain, costly to run, and impossible to trust in production pipelines.

The solution lies in classic software architecture principles: skills teach an agent what it knows, while commands give it something it can do. By decoupling domain judgment from deterministic execution, we build modular, testable, and cost-effective AI systems. A deterministic data pipeline command executes predictable steps automatically, invoking an AI skill only when genuine reasoning is required—such as when an unfamiliar payload arrives with an unmapped schema that threatens downstream tables.

Beyond the Persona: Mastering Agent Skills and Commands

In this post, we explore the architectural patterns behind modular agents, moving from an overburdened monolithic persona to a decoupled system that integrates human-in-the-loop security gates, Model Context Protocol (MCP) tools, and automated schema drift resolution.

🚀 Featured Open Source Projects

Explore these curated resources to level up your engineering skills. If you find them helpful, a ⭐️ is much appreciated!

🤖 AI Agents (ADK)

Focus: LLM Patterns, Skills & Commands, and Agentic Workflows
Status Topic

🏗️ Data Engineering

Focus: Real-world ETL & MTA Turnstile Data
Maintained License

📖 Data Engineering Process Fundamentals

Focus: Architectural patterns, hands-on data pipelines, and MTA turnstile case studies
Format

💡 Contribute: Found a bug or have a suggestion? Open an issue and be part of the open source project.

🔗 Related Repository: AI Engineering

Explore the full implementation of the skills, commands, and agent definitions used in this workflow:
https://github.com/ozkary/ai-engineering/tree/main/adk

YouTube Video

👍 Subscribe to the channel to get notified on new events!

📅 Agenda


The Monolith Agent: The Anti-Pattern of Context Bloat

When engineers first build AI agents, the standard approach is straightforward: open a prompt file or system message configuration and draft a comprehensive persona. This single file frequently includes:

  1. Agent Identity: Persona definition, voice, and conversational directives.
  2. Domain Business Logic: Field definitions, schema contracts, validation rules, and business nuances.
  3. Target System Rules: Naming conventions, SQL dialect specifics, partitioning strategies, and target warehouse instructions.
  4. Tool Definitions & Credentials: Instructions on which APIs to invoke and how to format payloads.

While this works for simple prototypes, in an enterprise data pipeline it becomes a severe anti-pattern known as the Monolithic Agent.

graph TD
    subgraph Monolithic Agent Anti-Pattern
        A[Single System Prompt] --> B[Domain Knowledge: MTA / Telemetry]
        A --> C[Dialect Rules: BigQuery / Snowflake]
        A --> D[Execution Logic: Table DDL / Storage]
        A --> E[Governance & Security Constraints]
    end
    B & C & D & E --> F[Context Bloat & Token Burn]
    F --> G[Hallucinations & Pipeline Crashes]

The Pitfalls of the Monolith

To solve this, we decouple the agent into Skills (declarative knowledge) and Commands (deterministic actions).


Demystifying Agent Skills: Teaching Domain Judgment (What It Knows)

A Skill represents domain expertise and semantic judgment. It answers the question: What does the agent know?

In practice, a skill is a structured specification (typically a modular Markdown manifest) that defines the rules, schema specifications, and analytical judgment required for a specific business domain.

graph TD
    subgraph Skills Architecture
        A[Agent Kernel]
        A -->|Injected via DI| B[MTA Transit Skill]
        A -->|Injected via DI| C[Factory Telemetry Skill]
        A -->|Injected via DI| D[BigQuery Governance Skill]
        B & C & D --> E[Domain Judgment & DDL Generation]
        E -.->|No Direct Execution| F[Structured Metadata Artifact]
    end

Core Characteristics of an Agent Skill

# Example: Injecting targeted skills into a lightweight SkillAgent
agent = SkillAgent(
    skills=[
        MtaTransitSkill(),        # Domain-specific schema & validation rules
        BigQueryGovernanceSkill() # Target warehouse conventions & partitioning rules
    ]
)
Domain Skills vs. Platform Governance Skills: A domain skill (like MtaTransitSkill) defines business entities, field requirements, and expected schema contracts. A platform governance skill (like BigQueryGovernanceSkill) defines engine-specific formatting standards—such as snake_case column conventions, timestamp parsing constraints, and lakehouse partitioning rules. Neither skill executes database queries; they purely supply the declarative rules the agent needs to synthesize compliant DDL.

By isolating domain and platform rules into distinct skills, teams can refine schema requirements and governance guidelines without modifying the underlying agent execution engine.


Demystifying Agent Commands: Deterministic Execution (What It Does)

A Command represents deterministic execution. It answers the question: What can the agent do?

In production data engineering, data warehouse modifications, bucket operations, and network transactions must never be left to stochastic LLM text generation. If you ask an LLM to generate and run an ALTER TABLE statement directly five times, you risk getting five subtly different variations—or worse, a hallucinated column that breaks downstream dashboards.

Commands provide absolute predictability by wrapping operations in testable, deterministic code that implements the Strategy Pattern.

graph TD
    subgraph Commands Architecture
        CMD[Agent Command] --> STRAT[WarehouseTableStrategy]
        STRAT --> BQ[BigQuery Strategy]
        STRAT --> SF[Snowflake Strategy]
        STRAT --> SQL[SQL Server Strategy]
        BQ --> MCP[Model Context Protocol MCP]
        SF --> MCP
        SQL --> MCP
    end

Why Commands Matter

  1. Guaranteed Determinism: The command implements fixed logic (e.g., GetFileSampleCommand, CreateExternalTableCommand). The LLM does not improvise the execution path.
  2. Strategy Pattern for Portability: The agent interacts with an abstract WarehouseTableStrategy. Whether the target is BigQuery, Snowflake, or Databricks, the agent invokes the same interface while the underlying strategy handles engine-specific semantics.
  3. Integration with Model Context Protocol (MCP): Commands communicate with target resources through standardized MCP servers, ensuring authentication, audit logging, and connection pooling are maintained centrally.
  4. Deterministic Routing via Configuration: Decisions about which data warehouse receives which feed are governed by deterministic configuration files (e.g., routing.yaml), rather than probabilistic prompt suggestions:
# routing.yaml: Deterministic ingestion targets
domains:
  mta_transit:
    target_warehouse: "bigquery"
    dataset: "transit_staging"
    strategy: "external_table_v2"
  factory_telemetry:
    target_warehouse: "snowflake"
    database: "iot_telemetry"

Managing Schema Drift: When Skills and Commands Collaborate

To understand how skills and commands interact in production, consider a real-world scenario from the Metropolitan Transportation Authority (MTA) New York City transit turnstile data—the canonical dataset explored throughout my book, Data Engineering Process Fundamentals.

The Problem: Incoming Schema Drift

In a modern Zero-ETL lakehouse, raw transit feeds land in cloud storage buckets as CSV or Parquet files. BigQuery queries these files directly via external tables, avoiding heavy ETL ingestion jobs.

When a new file format arrives (Version 2) containing altered column names, restructured timestamps, or additional device counters, traditional automated pipelines fail catastrophically. Unmapped columns crash downstream queries, while naive auto-ingestion can pollute warehouse tables with broken types.

sequenceDiagram
    autonumber
    actor Lake as Data Lake (Storage Trigger)
    participant Agent as SkillAgent Core (Skills & Commands)
    participant Gate as Security Gate (Pre-Hook & HITL)
    participant BQ as Data Warehouse (BigQuery via MCP)

    Lake->>Agent: Event: New File Dropped (turnstile_v2.csv)
    Note over Agent: 1. GetFileSampleCommand fetches 5-row sample<br/>2. MTA Skill detects schema drift<br/>3. Skill synthesizes normalized V2 DDL
    Agent->>Gate: Dispatch CreateExternalTableCommand(DDL)
    Note over Gate: Pre-Use Hook intercepts tool call<br/>Agent placed on hold (Suspended)
    Gate->>Gate: Security Prompt: Approve DDL Mutation?
    alt Human Approves (yes)
        Gate->>BQ: Execute DDL via BigQuery MCP Tool
        BQ-->>Agent: External Table Created (ext_mta_turnstile_v2 live)
        Note over Agent,BQ: Lakehouse reads V2 immediately at wire speed
    else Human Rejects (no)
        Gate-->>Agent: Mutation Aborted (Workflow Ends Safely)
        Note over BQ: Fail-Closed: Table is Never Created
    end

Deterministic Inspection & Schema Normalization

A critical design requirement of this architecture is deterministic inspection:

The Human-in-the-Loop (HITL) Gate

While the skill has the domain intelligence to normalize schemas, autonomous AI must never possess unchecked authority to modify production database structures.

To enforce security:

  1. Pre-Use Hooks Intercept Mutations: When the agent attempts to run CreateExternalTableCommand, the command dispatches an MCP tool call. The SecureToolAgent layer evaluates this call against registered Pre-Use Hooks. Read-only queries pass freely, but state-mutating operations (DDL/DML) are intercepted before they ever reach the database server.
  2. Agent Put On Hold: The pre-use hook suspends the agent's execution. The agent is placed on hold in a waiting state while a structured alert is surfaced to the human operator with the proposed DDL and target metadata.
  3. Approval (yes): If the human verifies the DDL and approves, the hook unblocks execution. The command forwards the query through the BigQuery MCP tool, provisioning the new table seamlessly.
  4. Rejection (no): If the human rejects the change or detects an anomaly, the process halts immediately and the table is never created. This fail-closed model guarantees that hallucinations, schema corruption, or unauthorized alterations cannot reach the analytical warehouse.

Live Demonstration: From Schema Drift to Warehouse Provisioning

During the live session, we walked through this pattern using Google Antigravity and the Agent Development Kit (ADK) CLI, operating on the open-source repository.

1. The SkillAgent Implementation

The agent class inherits from SecureToolAgent, which equips it with standard MCP tool connectivity and security controls. We then inject domain skills and executable commands:

class SkillAgent(SecureToolAgent):
    """
    Modular Agent demonstrating dynamic skill injection
    and deterministic command execution.
    """
    def __init__(self, skills=None, commands=None, routing_config="routing.yaml"):
        super().__init__()
        self.skills = skills or []
        self.commands = commands or []
        self.routing_config = load_routing(routing_config)

    def register_skill(self, skill):
        self.skills.append(skill)
        # Mounts domain specifications into agent knowledge boundary
        self.mount_knowledge(skill.get_manifest())

    def register_command(self, command):
        self.commands.append(command)
        # Binds executable tool signatures to the agent runtime
        self.bind_tool(command.get_tool_signature())

2. Simulating Schema Ingestion & The ADK CLI

To trigger the agent workflow, we invoke the Agent Development Kit (ADK) CLI:

adk agent run skill-agent \
  --event "storage.object.created" \
  --bucket "mta-turnstile-lake" \
  --path "feeds/v2/turnstile_2026_09.csv"

Understanding the ADK CLI Parameters

Production vs. Local Emulation: In a live enterprise deployment, developers do not manually trigger terminal commands. Instead, automated cloud infrastructure listens for file drops: a Cloud Storage bucket event publishes to Google Cloud Pub/Sub or Eventarc, which invokes a serverless container (such as Google Cloud Run or a Cloud Function) running the ADK agent runtime. The CLI command faithfully replicates this exact production event-driven behavior for local testing, CI/CD verification, and interactive playbook orchestration with Google Antigravity.

Deterministic Inspection in Action

Upon receiving the event, the agent executes GetFileSampleCommand. Rather than streaming entire megabytes into the prompt, it fetches a bounded 5-row sample. The MTA domain skill evaluates the sample against known schemas, identifies the drift, and synthesizes a normalized DDL statement:

[INFO] Executing: GetFileSampleCommand(uri="gs://mta-turnstile-lake/feeds/v2/turnstile_2026_09.csv")
[INFO] Sample loaded: 5 rows retrieved.
[WARN] Schema Drift Detected!
       Missing Baseline Columns: [C/A, UNIT, SCP]
       Detected V2 Columns:      [control_area, remote_unit, sub_device_id, net_entries, net_exits]
[INFO] Domain Skill generated BigQuery External Table DDL:
       CREATE OR REPLACE EXTERNAL TABLE `transit_staging.ext_mta_turnstile_v2`
       (
           control_area STRING,
           remote_unit STRING,
           sub_device_id STRING,
           event_timestamp TIMESTAMP,
           net_entries INT64,
           net_exits INT64
       )
       OPTIONS (
           format = 'CSV',
           uris = ['gs://mta-turnstile-lake/feeds/v2/*.csv'],
           skip_leading_rows = 1
       );

By normalizing the new external table schema to match lakehouse partitioning and naming conventions, the skill guarantees that the pipeline can consume the new file format without failing or interrupting existing workloads.

3. Human Approval (HITL) and Execution

Before mutating the data warehouse, the agent initiates table provisioning via CreateExternalTableCommand. However, because the agent inherits from SecureToolAgent, the tool call is intercepted by a Pre-Use Hook.

The Agent is Placed on Hold

The hook determines that creating or replacing an external table constitutes an infrastructure mutation requiring human authorization. The agent's execution is immediately suspended (put on hold), and an interactive prompt is rendered:

[SECURITY GATE] Action requires administrative authorization:
Mutation: Create BigQuery External Table 'ext_mta_turnstile_v2'
Do you approve the execution of this DDL? (yes/no): yes

Two Distinct Security Outcomes:

4. Beyond the Demo: Multi-Version Coexistence & Skill Evolution

While the live demonstration successfully created the V2 external table, provisioning a new schema is not the end of the architectural story—it represents the beginning of an ongoing, production-grade data lifecycle:

Multi-Version Coexistence in Production

With ext_mta_turnstile_v2 successfully deployed, the data warehouse operates in a live multi-version coexistence state. Downstream analytics, scheduled dashboards, and legacy ETL pipelines reading the original V1 table continue running without downtime. New applications and modernized reports can simultaneously begin querying the V2 schema. Over time, as data producers migrate and incoming V1 storage events cease, the legacy V1 external table can be gracefully deprecated and retired without risking breaking changes.

Continuous Evolution: The Inevitable Arrival of V3

Data feeds continuously evolve. Over time, an upstream vendor or sensor system might drop an unannounced Version 3 (v3) payload into the lakehouse. This arrival triggers the exact same event-driven loop:

  1. The storage trigger detects turnstile_2026_v3.csv.
  2. The agent executes GetFileSampleCommand to inspect the sample deterministically.
  3. The domain skill attempts to normalize the schema and synthesize an external table DDL.

When Complex Drift Causes the Agent to Fail

Not all schema changes are simple column additions. In real-world enterprise environments, a V3 payload might introduce complex edge cases—such as deeply nested JSON structures, conflicting timestamp formats, or ambiguous domain concepts that defy existing heuristic rules.

In these scenarios, the LLM-driven skill may fail to produce an acceptable or compliant schema proposal. It might hallucinate a data type, mishandle a nested field, or violate lakehouse partitioning rules.

The HITL Gate as a Circuit Breaker for Skill Updates

This failure mode is precisely why the Human-in-the-Loop (HITL) gate is indispensable:

Skills and Commands as First-Class SDLC Artifacts: Creating new skills or updating existing skills and commands is not a sign of pipeline failure—it is the natural, expected lifecycle of an agentic data system. Treating skills as version-controlled specifications ensures that as data domains evolve, agent intelligence grows incrementally alongside the business.

Key Takeaways & The Road Ahead

Refactoring from a monolithic agent to modular skills and commands bridges the gap between fragile AI experiments and reliable production data systems.

Concept Monolithic Agent Decoupled Skills & Commands
Domain Knowledge Crammed into one prompt Modular Markdown manifests (Skills)
Execution Stochastic LLM query generation Reusable, strategy-backed classes (Commands)
Token Cost High (full prompt on every call) Low (load only active domain and tools)
Portability Hardcoded to one database Abstracted across engines (BigQuery, Snowflake)
Safety & Governance Uncontrolled side-effects Pre-Tool Hooks & Human-in-the-Loop gates

Guiding Principles for Your Next Agent

  1. Skills Define What the Agent Knows: Keep them declarative. Use skills to codify domain heuristics, validation rules, and schema contracts.
  2. Commands Define What the Agent Does: Keep them deterministic. Use proven design patterns (Strategy, Factory) and route execution through Model Context Protocol (MCP) tools.
  3. Keep the Kernel Agnostic: Build a lightweight core engine and inject capabilities dynamically using Dependency Injection.
  4. Guard Critical Operations: Never allow an autonomous agent to execute destructive DDL or DML without human review. Use pre-tool hooks to enforce governance.

By establishing these boundaries, data teams can confidently harness generative AI to handle real-world data entropy while keeping pipelines robust, secure, and predictable.


🌟 Let's Connect & Build Together

Thanks for reading! 😊 If you found these architectural patterns helpful, let's connect and continue the conversation:

👉 Originally published at ozkary.com