Principal Platform & SRE Engineer — AI Infrastructure

I build the infrastructure that makes AI systems dependable in production.

25+ years running production platforms — from owning release and reliability engineering for Hilton's global rebuild (monthly releases to one or more per day, as needed, at 99.9%+ availability) — now pointed at the newest class of production system: AI. Not model training. The systems around the models — provider routing and fallback, accuracy validation of LLM output, cost governance, and guardrails for agentic systems.

The proof is public: Seeker OS generates documents under claim-level accuracy enforcement, forge took a real container image from 507 CVEs to 0 behind a deterministic CI gate, and telemetry-gcp went from zero GCP experience to live on Cloud Run in a single day. Hands-on IC by preference: 80% building, 20% mentoring — and I use AI as a force multiplier, with three decades of judgment steering it.

Featured Project

Seeker OS

Seeker OS screenshot

An AI-powered job search pipeline whose LLM output is claim-level accuracy-enforced — the system is structurally incapable of lying on my behalf.

Stack
AI Infrastructure FastAPI LLM

Latest Writeup

View all

From Chat Agent to Pipeline: Building SeekerOS

The first post in a series on running LLMs in systems that have to be correct.

Read

Selected Experience

View more on LinkedIn
  • Sep 2024 - Present
    Zapcom Group
    Principal Engineer

    Clients: Mass General Brigham · Accelya · Remote

    • Building an Internal Developer Portal (IDP) for Mass General Brigham — self-service provisioning of VMs and storage for clinical researchers, golden path architecture extensible to additional resource types.
    • Designed and standardized GitLab CI pipelines with reusable templates across client engagements, accelerating release cadence and reducing pipeline maintenance burden.
    • Built an AI-assisted AWS Security Group optimization tool after tracing recurring access failures to years of accreted manual rules — used VPC flow logs to derive least-permissive rule sets, eliminating the bulk of manual security operations overhead.
    • Wrote onboarding documentation that ramped client engineers onto the pipeline architecture.
    • (Accelya) Rewrote and hardened legacy PowerShell across multiple GitLab CI pipelines, resolving recurring reliability failures and stabilizing build and release automation.
    • (Accelya) Built repository-decoupling pipelines that extracted individual airline codebases from a monolithic repository into dedicated per-airline repos and initialized them for independent development.
    • (Accelya) Developed dependency-build and artifact-publishing pipelines that delivered per-airline artifacts to a shared location for deployment by the client’s existing GoCD/Spinnaker platforms.
    • (Accelya) Streamlined multi-pipeline navigation by building a GitLab Pages landing page that deep-linked into the native GitLab pipeline UI.
    GitLab CI AWS Azure DevOps
  • Nov 2016 - Jun 2020
    Hilton Worldwide
    Director, Digital Release & Reliability Engineering

    Collierville, TN

    • Owned release management and SRE for Hilton’s web and mobile app rebuild — driving the engineering, process, and cultural changes that took the organization from monthly releases to one or more per day, as needed — release-on-demand in practice. Maintained 99.9%+ availability across web and mobile properties serving millions of guests globally throughout the transition.
    • Led a 23-member Release & Reliability Engineering team of top performers to support the demands of Hilton’s cloud-native transformation.
    • Managed $5M+ vendor relationships with Akamai and Dynatrace, driving strategic alignment and delivering measurable platform improvements.
    • Championed “You Build It, You Run It” engineering culture — embedding dev ownership of production reliability through architecture decisions and influence. Teams that own their 2am incidents write fundamentally better code.
    • Launched a dedicated performance engineering function from scratch — establishing the discipline, tooling, and practices that supported software engineering teams across the platform.
    • Delivered $400K in annual infrastructure savings through Akamai Image Manager optimization.
    • Blocked 90%+ of malicious bot traffic via Akamai Bot Manager, protecting platform integrity and reducing fraudulent activity.
    • Integrated Dynatrace APM across the platform, accelerating incident detection and resolution by 20%.
    • Managed JFrog Artifactory as the enterprise artifact repository — dependency governance, reproducible builds, and secure artifact delivery across all engineering teams.
    SRE Akamai Dynatrace JFrog Artifactory

Projects

View all

Let's connect.

If you want to get in touch about something or just to say hi, reach out on social media or send me an email.