Sep 2025 — Present
Spin
Mexico · Remote
Sr Observability Infrastructure EngineerIC4Present
I scale system visibility across the company. I co-manage multiple Datadog tenants — a critical platform used by 1,200+ people for logs, metrics, traces and events — and the strategy behind 6,000+ production monitors across the business units, making sure critical alerts come first. My focus is turning technical observability into business impact: custom metrics centered on the customer-facing operation, operational dashboards at infrastructure and business level, and escalation and routing policies integrated with Jira Service Management (ITSM).
Key achievements
- Multi-tenant alerting management for 1,200+ users and 6,000+ monitors.
- 60% reduction in false-positive alerts through systematic monitor calibration.
- Service instrumentation with OpenTelemetry and a custom span attribute strategy, the foundation for custom business metrics monitored in real time and operational dashboards.
- Alerting Governance as Code with Terraform: centralized routing rules, schedules and escalation policies for 50+ teams across 5 business units, federating roles and eliminating manual platform configuration.
- Optimized escalation and routing in Jira Service Management, categorizing alerts by Golden Signals (requests, errors, latency) to align response with real user impact.
- Built an AI agent on AWS Bedrock, currently in beta, that attends Datadog alerts, runs triage and detects impact on the operation, prioritizing what affects end customers.
- Agent harness and Spec Driven Development (SDD) across the company's repositories for agent orchestration: versioned rules and specs that everyone using Claude in a repository follows alike, improving developer experience.
Stack
- Datadog
- OpenTelemetry
- AWS Bedrock
- Claude
- Terraform
- Jira Service Management
- CloudWatch
- Prometheus