Sobre a oportunidade
About the role We are looking for a DevOps Engineer to manage application migrations between environments and maintain production and staging system reliability. This person handles incident response and on-call support, while improving observability, automation, and Kubernetes-based infrastructure. Weekend availability for migrations and comfort with CI/CD pipelines and SLAs are essential. What you will do Migrate applications between environments; these migrations typically take place on weekends, so weekend availability is required. Monitor and support production and staging environments in real time, ensuring high availability, performance, and stability. Respond to incidents, perform triage and root cause analysis, and contribute to post-incident reviews and remediation efforts. Participate in an on-call rotation with defined SLAs. Handle ad-hoc and unplanned operational requests from Product, Support, and internal teams. Maintain and enhance monitoring, alerting, dashboards, logs, and metrics; improve signal-to-noise ratio and standardize observability practices. Support CI/CD pipelines, production releases, and GitOps workflows. Contribute to automation efforts to reduce operational toil. Maintain and improve Kubernetes-based infrastructure and containerized workloads. Support Infrastructure as Code practices and ongoing environment improvements. Must haves At least 2 years of experience in Site Reliability Engineering, DevOps, or Production Operations. Demonstrable AWS experience supporting production environments. Experience supporting production SaaS applications. Strong understanding of CI/CD systems (GitHub Actions, Jenkins, CircleCI, or similar). GitOps experience and strong Git fundamentals. Experience using GitHub, Jira, and Confluence in collaborative engineering environments. Kubernetes experience (EKS, kOps, or similar). Docker/containerization experience. Observability stack experience (Grafana, Prometheus, Loki, PagerDuty, or similar). Scripting experience (Bash, Python, or Go). Infrastructure as Code experience (Terraform, Helm, or similar). Working knowledge of relational databases (e.g., PostgreSQL, MySQL) for troubleshooting and operational support. Comfortable working within structured operational processes and SLAs. Strong written and verbal English communication skills; able to clearly explain technical concepts. Self-driven with a growth mindset. Weekend availability for application migrations between environments when needed. Nice to haves AWS certifications (Solutions Architect, DevOps Engineer, SysOps Administrator, etc.). Experience in multi-tenant SaaS environments. Experience working in globally distributed teams. Familiarity with ChatOps practices. Experience improving monitoring quality and reducing alert fatigue. Background in operational cost optimization. Perks and benefits Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized Well-being & support: access local well-being programs and people-focused support tailored to your location Job Type: Full-time Work Location: Remote