Job Description

Overview Position Summary: With team members and customers in 39 countries around the globe, HostPapa is one of the fastest-growing web hosting companies offering a wide range of products. We provide tools and services critical to online success, including a Website Builder, award-winning customer support, email, and cloud-based solutions that put customers first. This role focuses on CloudBlue, a HostPapa business that powers cloud commerce for many of the worlds largest service providers, including major Telcos, distributors, and MSPs. CloudBlue enables partners to monetize and manage cloud services and subscriptions at scale, combining the agility of a high-growth business with the backing of a global organization. As the Site Reliability Engineer, you will help ensure the reliability, scalability, and observability of CloudBlues multi-tenant SaaS platforms used by service providers worldwide. You will focus on improving system stability and performance through monitoring, high availability, and incident response, while working closely with DevOps, Platform, and Engineering teams to build and operate resilient production systems. What youll do Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services to ensure reliability and performance Influence system architecture with a strong focus on reliability, scalability, and operability, designing systems for fault tolerance, graceful degradation, and self-healing Reduce operational toil by identifying opportunities for automation and process improvement Design and operate CloudBlues observability stack across metrics, logs, and traces using tools such as Datadog, Grafana, and Elastic Stack Develop actionable alerting strategies and dashboards that provide clear insight into platform and business health Design and maintain high-availability architectures, implementing redundancy, failover, and disaster recovery strategies across regions and availability zones Conduct capacity planning, load testing, and performance optimization to ensure platform stability and scalability Act as a senior responder during production incidents, leading incident coordination, communication, and service restoration Own blameless postmortems and drive improvements that reduce incident frequency, MTTR, and customer impact Improve reliability of Kubernetes-based platforms through health checks, autoscaling strategies, rollout safety, and resilience testing Partner with engineering and DevOps teams to improve deployment safety, rollback strategies, and platform reliability Maintain runbooks and operational documentation, and promote SRE best practices across engineering teams Support other tasks or projects as assigned to meet team and business needs About you 3+ years of experience as an SRE, DevOps Engineer, or Production Engineer, with strong ownership of production systems Proven experience operating highly available, enterprise-grade, multi-tenant SaaS platforms Hands-on experience with observability and monitoring tools such as Datadog, Grafana, and Elasticsearch/Kibana Solid understanding of Linux, networking, and distributed systems fundamentals Experience working with containerized environments such as Docker and Kubernetes Strong scripting and automation skills using Python and/or Bash Experience participating in on-call rotations and incident response in production environments Strong written and spoken English Experience defining SLIs/SLOs and managing error budgets at scale will be considered a plus Exposure to hyperscale or service-provider-grade platforms is an advantage Cloud experience, preferably with Azure; experience with AWS and/or GCP will also be valued Experience working with hybrid or on-premises integrations is beneficial Familiarity with chaos engineering and resilience testing will be considered an asset What We Offer Work from anywhere - this is a remote opportunity A competitive salary that values you and your unique skill sets Career advancement & professional development opportunities to help you reach your full potential Flexible work arrangements to support work/life balance About Us At HostPapa, weve been committed to providing a complete array of enterprise-grade cloud services solutions to every business owner since 2006. These services are offered in a one-stop shop, with 24/7 award-winning customer support in four languages. Our HostPapa team values diversity and inclusion. We have a friendly company culture built on trust and respect. With the acquisition of several companies into our product portfolio, were growing at an incredible rate and have ample opportunities for career growth. Come join our talented team of enthusiastic, hard-working, passionate, driven people engaged in meaningful, innovative work. We cant wait to meet you! HostPapa is an equal-opportunity employer committed to diversity and inclusion. We encourage individual achievement and recognize the strength of our diverse team. HostPapa is committed to providing accommodations for people with disabilities. If you require accommodation, please let us know, and we will work with you to meet your needs. Accommodation may be provided in all parts of the hiring process. It is anticipated that this position will be performed outside of Ontario. #J-18808-Ljbffr

Job Title

Company : Host Papa Inc

Location : Toronto, Ontario

Created : 2026-03-21

Job Type : Full Time