Twilio Software Engineer, Platform Engineering (L2) at Twilio responsible for developing and managing large-scale distributed systems. Works on Kubernetes, cloud infrastructure, and automation to ensure system reliability and scalability.
Responsibilities
WEAR THE CUSTOMER’S SHOES: Design, build, and operate services and automations to manage kubernetes clusters at scale. Partner closely with product management and technical leadership to break down complex system requirements into manageable, iterative milestones.
BE AN OWNER (CODE QUALITY): Drive rigorous code reviews and push for maintainable patterns in our codebase, ensuring high testing standards (unit, integration, and component testing) are executed across the team and platform.
BE AN OWNER (INFRASTRUCTURE & OBSERVABILITY): Manage and enhance cloud configurations across AWS and Azure environments utilizing Infrastructure as Code (Terraform). Ensure deep observability coverage by standardizing metrics, alerts, and distributed tracing across core data pipelines.
CHAMPION ENGINEERING HEALTH: Advocate for a clean architectural foundation. Proactively identify technical debt, system bottlenecks, and single points of failure (SPOF), balancing feature delivery with critical platform refactoring.
MENTOR AND LEAD: Foster a collaborative environment by mentoring junior engineers, leading technical sprint planning, and sharing expertise across distributed engineering nodes.
Qualification
ExperienceModern Development WorkflowDeployment OrchestrationLanguage ProficiencyCloud InfrastructureInfrastructure as CodeDistributed SystemsSystems Mindset
Required
Twilio values diverse experiences from all kinds of industries, and we encourage everyone who meets the required qualifications to apply. If your career is just starting or hasn't followed a traditional path, don't let that stop you from considering Twilio. We are always looking for people who will bring something new to the table!
Experience: 4+ years of professional software engineering experience building and operating resilient backend services at scale using Kubernetes. Experience with CAPI, EKS and managing zero-downtime Kubernetes cluster upgrades, including node draining, API deprecations, and PodDisruptionBudgets.
Modern Development Workflow: Practical experience leveraging AI-assisted development tools (e.g., Claude Code) to accelerate code generation, automate testing, and streamline debugging workflows or strong desire to learn.
Deployment Orchestration: Hands-on Experience implementing GitOps workflows with ArgoCD and automated pipeline orchestration with Harness (or an equivalent enterprise CI/CD platform).
Language Proficiency: Strong, hands-on experience with Shell, Terraform, Yaml and Go (Golang).
Cloud Infrastructure: Solid experience deploying and managing production workloads in cloud environments - ideally with deep exposure to AWS core services (such as EKS, EC2, S3) or their Microsoft Azure equivalents (such as AKS, Virtual Machines, Blob Storage). Understanding of container networking (VPC/VNet, pod IPAM, CNI plugins.
Infrastructure as Code: Proficiency with Terraform for provision-level automation and maintaining environment parity.
Distributed Systems: Strong theoretical and practical understanding of distributed datastores, caching layers, and asynchronous event streaming (e.g., Kafka or similar queuing ecosystems).
Systems Mindset: Strong foundational background in computer science fundamentals, data structures, and building self-healing cloud architectures.