Tinybird
Site Reliability Engineer
Tinybird
€58k - €97k
Spain / Remote
Kubernetes
AWS
GCP

Site Reliability Engineer

Overview

At Tinybird, we help developers and data teams unlock the power of real-time data.

Job Description

Tinybird is the essential tool that data engineers and software developers have been waiting for, making it easier to drive innovation.

Responsibilities

  • - Enhancing high availability and elasticity so the system can scale automatically and efficiently as our customer base grows.
  • - Boosting our observability capabilities, from low-level resource usage to high-level service metrics, including telemetry, dashboards, alerting, and long-term visibility into system health.
  • - Improving disaster recovery with better tools, incident discovery, and enhanced oncall experiences.
  • - Handling Kubernetes lifecycle tasks, managing cluster infrastructure, autoscaling, and ensuring safe deployments.
  • - Understanding how ClickHouse operates under the hood and extracting the best performance possible from it.
  • - Identifying bottlenecks and improving performance across storage, networking, and compute.
  • - Reducing operational burden by transforming manual or fragile processes into repeatable, well-managed systems.
  • - Assisting with incident prevention, operational reviews, and follow-up tasks after reliability issues.
  • - Strengthening CI/CD foundations to help teams build and deploy changes with greater confidence.

Required Skills

  • - Strong experience in designing, building, and running distributed cloud architectures and large-scale web-based production systems.
  • - Deep knowledge of Kubernetes, which is essential for this role.
  • - Comfortable designing and operating production-grade clusters, writing custom controllers or operators as needed, and tuning autoscaling mechanisms (KEDA, Karpenter, and similar) to respond to real-time workloads.
  • - Skilled in AWS and GCP.
  • - Coding skills are required, primarily in Python and some C++.
  • - Comfortable operating close to production: debugging incidents, understanding system behavior, improving observability, and enhancing service reliability.
  • - Think in systems and pay attention to edge cases, failure modes, and specific implementation details.
  • - Care about performance, reliability, cost efficiency, and operational simplicity.
  • - Communicate clearly in writing and use AI tools to enhance efficiency and improve workflows.
  • - Fluent in English and Spanish.

Benefits

  • - 22 days of holiday a year (plus your birthday and public holidays).
  • - Freedom to work from wherever suits you best.
  • - Up to €2,800 to help you set up your home workspace.

About the company

Tinybird helps data teams build real-time Data Products at scale through SQL-based API endpoints


All Job Openings at Tinybird