Datadog Illuminated: Evaluating an AI-Powered Process That Creates and Tunes Monitors, with Eric Karschner

LLMday Austin: The Best Laid Spans: An AI Agent That Instruments Your Code with OpenTelemetry

We secure data at rest. We secure data in transit. But what about data in use?

In this episode of 🌩️Thunder, Tobin Feldman-Fitzthum explains how Confidential Containers uses hardware-level trusted execution environments to isolate your Kubernetes workloads from everything outside them, including the Kubernetes control plane and the cloud provider itself.

Have a look-see! youtu.be/WdW-_KU_B…

Whitney Lee on the left, laughing, wearing a dark jacket with bangs. Tobin Feldman-Fitzthum on the right, shown in a black-and-white photo. Large yellow text reads Confidential Containers and white text below reads Hardware Protection for Kubernetes Workloads, with the name Tobin Feldman-Fitzthum at the bottom in yellow. A lightboard covered in handwritten technical notes fills the background. The 🌩️Thunder show logo appears in the top-left corner.

Datadog Illuminated: Evaluating an Agent That Creates and Tunes Monitors, with Eric Karschner

kagent is a framework to help you make AI agents and run them in Kubernetes declaratively, with YAML.

Lin Sun, co-creator of kagent, joined to walk through how it works. Here’s the lightboard we made together.

Read the notes: gist.github.com/wiggitywh…

A hand-drawn lightboard diagram on a black background, filled with colorful marker text and boxes mapping out how kagent works. Sections include "Before kagent," "kagent config" with a Kubernetes CRD circle, "kagent controller watches," "Agent Sandbox," "Agent Runtimes," "Why run your agent on Kubernetes," "AUTH," and "How do people use kagent," connected by arrows and underlines in pink, yellow, and green.

Datadog Illuminated: Evaluating an AI-Powered Process That Creates and Tunes Monitors, with Eric Karschner

Will AI help or hinder a developer’s experience with an internal platform? At KubeCon EU, the audience found out. My bestie Viktor Farcic (Upbound) and I ran an interactive session where attendees voted at key moments on which tools our AI agent would use. I built the demo to work with any combination they chose, on the fly.

The agent investigated a broken app, used semantic search to navigate cluster resources, and deployed on a developer’s behalf, with no cluster access required. Expect surprises, stumbles, and a real conversation about what AI is (and isn’t) ready for in platform engineering.

Watch the recording here: youtu.be/k7ct4sW-9…

Title slide on an orange background. Large bold white text reads "Choose Your Own Adventure: AI Meets Internal Developer Platform." Below in white italic: "Whitney Lee, Datadog & Viktor Farcic, Upbound." The KubeCon and CloudNativeCon Europe 2026 logos appear in the upper left corner. A stylized windmill and Dutch tulips decorate the right side and bottom corners.

Software Defined Interviews: DevOpsDays, Technical Storytelling, and Loading the Dishwasher, with Jason Yee

Datadog Illuminated: Building an Eval Platform for an Investigation Agent, with Benjamin Barton

In this Datadog Illuminated episode, Scott Yak & I discuss how evals for the Datadog MCP server run against live data that goes stale FAST

How did Scott & team solve this?

See the lightboard below & read the board notes here: gist.github.com/wiggitywh…

Handwritten green, yellow, and white notes on a black lightboard titled "Evaluating the Datadog MCP Server." Left column shows "Why eval MCP server?" with a diagram of a query flowing from a person through an Agent (LLM), Datadog MCP, and Datadog Backend and back. Below it, "How eval MCP server?" defines eval = Agent harness + MCP server + Q&A pair, with Actual ≈ expected response (if passing). Middle column shows "Automate eval generation," a flow starting with a seed query fanning out into "FUZZING" to make many question variants, with a note that one seed produces many Q/A pairs and evals get re-run to update answers. Right side lists "Benefits of Great Evals" in four numbered sections: Speed (fast to generate, fast to run, see impact quickly), Visibility Into Progress (high quantity of evals, tagging, traces), Mutual Benefit (Datadog AI agents like Bits AI SRE and Bits AI Assistant use the Datadog MCP, so a better MCP server means better agents), and Devs Like Writing Evals (fast feedback, showing impact, preventing regression).

AI Engineer World’s Fair: Build a Platform, Unleash an Agent on it…. and Watch it Burn!

At Cloud Native AI Day (KubeCon EU), Thomas Vitale (Systematic) and I co-presented a live demo where audience votes drove real-time canary rollouts for GenAI apps using OpenTelemetry and Flagger.

Watch the recording here: youtu.be/CRcYbl-g3…

Title slide with a light blue gradient background. Large dark blue bold text reads: Rollout on Reception: Progressive Delivery for GenAI Apps Using Real-Time User Feedback. Below in dark blue italic: Thomas Vitale, Systemic and Whitney Lee, Data Dog. In the upper right corner: the Cloud Native AI + Kubeflow Day Europe logo, a hexagonal badge with a brain icon and EUROPE in a blue box.

AI Engineer World’s Fair: Build a Platform, Unleash an Agent on it…. and Watch it Burn!

Software Defined Interviews: DevOpsDays, Technical Storytelling, and Loading the Dishwasher, with Jason Yee

Software Defined Interviews: Running D&D with AI, Determinism First, and Workflow ROI, with Michael Rishi Forrester

Datadog Illuminated: Building an Eval Platform for an Investigation Agent, with Benjamin Barton

SREday Austin: Livin' In the Future: Your Platform’s Next Interface Is an AI Agent

Datadog Illuminated: How to Use AI to Find Security Flaws in Static Code, with Bahar Shah

This is the finished lightboard from my 🌩️Thunder conversation with Volkan Özçelik about SPIFFE and SPIRE.

Bottom line: SPIFFE is a secure and automated way to manage identity, no secrets managers needed.

How does it work? gist.github.com/wiggitywh…

A black lightboard covered in colorful handwritten notes organized into four columns. The leftmost column, headed "Before SPIFFE...", lists problems with service tokens and mTLS certificates and defines SPIFFE and SPIRE. The second column, headed "WHO + WHAT = POLICY," includes a small diagram of pods stacked above a host layer and a machine layer, illustrating where identity is best assigned. The third column, headed "SPIFFE is designed to solve these problems," lists SPIFFE IDs, SVIDs, and how the Workload API attests and issues identity to workloads. The rightmost column covers what happens once attestation completes, a boxed list of "Benefits of SPIFFE," and boxed definitions of SPIRE Server and SPIRE Agent.

Software Defined Interviews: The Chief Therapy Officer, Job Titles as Culture Change, and AI as Spreadsheets, with Bryan Ross