Skip to content
Seba Pinery

Career

Experience

A career focused on SRE, observability and infrastructure — from backend development to cloud engineering at scale.

Remote-first since 2021

  1. Sep 2025 — Present

    Spin

    Mexico · Remote

    Sr Observability Infrastructure EngineerIC4Present

    I scale system visibility across the company. I co-manage multiple Datadog tenants — a critical platform used by 1,200+ people for logs, metrics, traces and events — and the strategy behind 6,000+ production monitors across the business units, making sure critical alerts come first. My focus is turning technical observability into business impact: custom metrics centered on the customer-facing operation, operational dashboards at infrastructure and business level, and escalation and routing policies integrated with Jira Service Management (ITSM).

    Key achievements

    • Multi-tenant alerting management for 1,200+ users and 6,000+ monitors.
    • 60% reduction in false-positive alerts through systematic monitor calibration.
    • Service instrumentation with OpenTelemetry and a custom span attribute strategy, the foundation for custom business metrics monitored in real time and operational dashboards.
    • Alerting Governance as Code with Terraform: centralized routing rules, schedules and escalation policies for 50+ teams across 5 business units, federating roles and eliminating manual platform configuration.
    • Optimized escalation and routing in Jira Service Management, categorizing alerts by Golden Signals (requests, errors, latency) to align response with real user impact.
    • Built an AI agent on AWS Bedrock, currently in beta, that attends Datadog alerts, runs triage and detects impact on the operation, prioritizing what affects end customers.
    • Agent harness and Spec Driven Development (SDD) across the company's repositories for agent orchestration: versioned rules and specs that everyone using Claude in a repository follows alike, improving developer experience.

    Stack

    • Datadog
    • OpenTelemetry
    • AWS Bedrock
    • Claude
    • Terraform
    • Jira Service Management
    • CloudWatch
    • Prometheus
  2. Jan 2024 — Sep 2025

    Spin

    Mexico · Remote

    Site Reliability EngineerIC3

    I worked on critical incident response, coordinating cross-functional teams and applying ITIL practices to improve escalation, documentation and communication. With the Incident Managers I refined alert prioritization and post-mortem processes and wired monitoring into incident workflows. I led root-cause analyses of recurring issues and proposed architectural improvements — circuit breakers, rate limits, caching — to increase resilience, and helped embed SRE practices (service ownership, on-call rotations, reliability reviews) so product teams could own the health, latency and availability of their services.

    Key achievements

    • Critical incident response in an on-call rotation.
    • Alert monitors for both development and operations teams.
    • Post-mortems and root-cause analyses.
    • Tracking of recurring problems and architectural improvement proposals.
    • SLA tracking with service providers.

    Stack

    • AWS
    • Kubernetes
    • Terraform
    • ArgoCD
    • Argo Rollouts
    • Helm
    • Linkerd
    • Jira Service Management
    • Datadog
  3. Jun 2022 — Jan 2024

    Ventura Travel

    Germany · Remote

    Infrastructure Engineer

    I led the full rebuild of the Google Cloud Platform infrastructure, designing it from scratch and deploying it with Terraform: Kubernetes clusters, Cloud Storage and cloud networking (VPCs, subnets, firewall). The result was a markedly more stable and reliable platform and a solid, scalable foundation for production.

    Key achievements

    • 75% reduction in GCP infrastructure costs.
    • CI pipelines on Spot instances in GKE: 50% faster runs and twice as many daily deployments.
    • MTTA down 75% and MTTR down 50% after rolling out Prometheus, Loki and Grafana with alert coverage for 100% of microservices.
    • Infrastructure governance with Terraform and FinOps practices for cloud budget control.

    Stack

    • Google Cloud
    • Kubernetes
    • Terraform
    • Grafana
    • Prometheus
    • Loki
    • GitLab CI
    • ArgoCD
    • Traefik
    • Cloudflare
  4. Aug 2021 — Jun 2022

    Time Jobs

    Chile · Remote

    DevOps Engineer

    I took part in the strategic transformation of the company's cloud platform, executing its full migration from AWS to Google Cloud Platform. This ensured business continuity and laid the groundwork for a scalable, cost-efficient architecture through an end-to-end GitOps strategy: Terraform for provisioning and ArgoCD to automate microservice deployments, standardizing workflows across the organization.

    Key achievements

    • Complete AWS to GCP migration with no operational downtime.
    • Private libraries for custom OpenTelemetry spans and attributes, improving observability quality and speeding up the diagnosis of complex issues.
    • Declarative deployment strategy with Terraform and ArgoCD, unifying and automating the lifecycle of infrastructure and microservices across the organization.

    Stack

    • AWS
    • GCP
    • Terraform
    • ArgoCD
    • OpenTelemetry
    • Prometheus
    • Grafana
    • Loki
    • GitLab CI
    • Traefik
    • Cloudflare
  5. Until 2021

    Various companies

    Remote

    Backend Developer

    Backend development with Node.js and TypeScript before moving into infrastructure and reliability.

    Key achievements

    • Microservice development with Node.js, TypeScript and NestJS.

    Stack

    • Node.js
    • TypeScript
    • NestJS
    • MongoDB
    • PostgreSQL
    • Docker

Let's work together?

I'm available for consulting and collaborations. Let's talk about how I can help with your infrastructure.

Start a conversation