Site Reliability Engineer
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
OXIO is the first NeoTelco. We are building the world’s largest, most accessible, and insightful Telecom network. Our platform empowers anyone to spin up their own carrier from a browser, scaling and supporting you as you scale your network to millions of users.
We ensure that users and devices are connected, and stay connected wherever they go: cross-country, carrier, or cellular technology. We help them pay less for mobile data. This technology is provided through our Carrier-as-a-Service platform: BrandVNO, a fully customizable telecom service. In addition, we enable clients of our service to extract the value from telecom data—enriching their customer experience, business intelligence, and product understanding in the many markets in which we operate.
Come join us in creating a modern technology platform with a group of engineers dedicated to advancing our vision. Our team is passionate about what we build, open to new ideas and challenges, and has our sights set on the future of connectivity.
Responsibilities
Design and implement platform on the cloud to support OXIO backend services
Automate technical operations: deployments, scaling, recovery, etc.
Monitor and maintain mission-critical production infrastructure to ensure maximum uptime
Participate in an on-call rotation and culture of continuous improvement through blameless postmortems
Enable the Engineering/Telecom/Data Engineering teams by providing them the tools to operate the service they build
Essentials
Understanding of Linux/Unix systems (most systems are Linux-based)
Familiarity with Linux/Unix system internals like process management, filesystems, memory management, and networking
Proficiency in at least one programming language (Python, Go, or Ruby) and strong skills in scripting (Bash, Perl)
Experience with infrastructure provisioning tools such as Terraform, CloudFormation, or Ansible
Familiarity with containerization (Docker) and orchestration tools (Kubernetes)
Familiarity with monitoring tools like Prometheus, Grafana, or Datadog
Knowledge of setting up alerts, analyzing logs, and creating dashboards for observability
Familiarity with incident management practices (e.g., runbooks, postmortems)
Experience in being part of an on-call rotation and handling incidents
Experience in setting up and maintaining Continuous Integration/Continuous Delivery pipelines (Jenkins, GitLab CI, CircleCI, etc.)
Hands-on experience with cloud providers (AWS, Google Cloud, Azure)
Knowledge of virtualization technologies (VMware, KVM) and cloud-native architecture
Understanding of TCP/IP, DNS, HTTP/HTTPS, load balancing, and firewalls
Nice to Have
Strong understanding of deployment strategies (canary releases, blue-green deployments, etc.)
Familiarity with high availability and understanding failover mechanisms
Familiarity with IAM (Identity and Access Management) and zero trust principles
Experience working with distributed systems (e.g., Kafka, Cassandra, Elasticsearch)
Building custom monitoring tools or writing complex automation scripts
Functional knowledge of database management (SQL and NoSQL)
Familiarity with distributed tracing (Jaeger, OpenTelemetry) and advanced log aggregation strategies (ELK stack, Splunk)
Familiarity with performance profiling tools and optimizing application performance under heavy load
Familiarity in load testing and identifying bottlenecks
Familiarity with Configuration Management using SaltStack for maintaining server configurations
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
Solutions Engineer (API) - Healthcare SaaS
Verifiable
United StatesHead of Venture Creation - Deep-Tech Company Builder
Roadrunner Venture Studios
United StatesSr. Database Reliability Engineer II
NinjaTrader
United StatesSenior Identity & Access Architect (Agent-Native Systems) at Keycard
Keycard
United StatesExplore Top Companies in this Space
Soracom
Telecommunications / IoT / Cloud Computing / SaaS
fonio
Artificial Intelligence / SaaS / Telecommunications
Ooma
Telecommunications / Software / VoIP / AI
E-Space
Sustainable LEO Satellite Constellations / Smart-IoT & AIoT Aerospace Telecommunications / Sovereign Defense Networks & Space Debris Removal Tech
OXIO (operating at oxio.com) is a telecom-as-a-service (TaaS) platform engineered for brands, enterprises, and mobile virtual network operators (MVNOs). Founded in 2018 and headquartered in New York, OXIO redefines telecom infrastructure by unbundling traditional mobile network components, eliminating the need for costly, proprietary systems. Under the hood, the platform leverages modular, no-code tools that democratize network customization—allowing businesses to deploy, manage, and scale mobile networks without deep technical expertise. This empowers MVNOs, retailers, and global enterprises to launch their own mobile services with unprecedented flexibility, cost efficiency, and innovation speed. Recognized as a Forbes Top 50 Startup Employer of 2023 and named one of America’s Best Startup Employers, OXIO has also expanded its footprint to Mexico City and Montreal, catering to a broader market while fostering a culture of connectivity-driven disruption.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.