Back to Jobs
United States
Development 2h ago
Site Reliability Engineer (SRE)
United StatesFull-time
$131,500 — $175,500 USD
Mid-Level
Be the first applicant! 🚀
Job Description
Key Skills Required
Master these to land this role
DevOpsBestseller 🔥
Learn in 63 HoursPythonBestseller 🔥
Learn in 56 HoursCybersecurityAWSKubernetes
Want to know if you're a match for this job?
Your Role:
Have you heard of Tenable? Our cloud-based vulnerability management platform built for today’s dynamic IT assets, like cloud, containers and web apps? Well, that’s what you’ll be working on in this role. You will need to continue to quickly build out the platform, scale it automatically, and make it more self-managing for our private and US Government cloud customers!
Your Opportunity:
- Responsible for taking the code and functionality of Tenable cloud products and ensuring they’re reliable and highly available in cloud environments
- Responsible for responding to support escalations which involve troubleshooting complex technical problems and resolving data/configuration issues within defined service level objectives
- Responsible for developing software, tools, and scripts to automate deployment, management, and monitoring of production systems in all environments
- Collaborate with peers on complex projects
- Collaboration with cloud engineers in understanding new cloud technologies, assessing impact to security services operations, and proposing solutions to existing business problems
- Collaboration in the software development lifecycle to develop detailed enhancement/bug definitions, write functional requirements, translate the requirements into solution designs, and navigate the functional requirements through to Production deployments
- Proactively look for ways to create efficiencies within operations as it pertains to the tools and technology used by Tenable to support their customer base
- Manage, participate in, or directly work on any additional projects, assignments, or initiatives assigned by management
- Create/maintain documentation for operational procedures
- Contribute to standardization efforts across disciplines and services in conjunction with embedded SREs throughout the organization
- Document and perform system upgrades, application updates, and define monitoring requirements based on customer or organization needs
- Participate in an on-call rotation and support 24x7 availability of production application systems
- Remediate infrastructure and/or container security vulnerabilities
What You'll Need:
- U.S citizen required
- 2+ years of related SRE experience
- Apply core software engineering fundamentals* (e.g., data structures, algorithms, concurrency, system design) to solve complex reliability and scalability challenges at a distributed level
- To be considered for this role, you must meet one of the following criteria: Currently reside in the San Francisco Bay Area, CA and able to work remotely
- Experience with Terraform or similar IaC technologies
- A curiosity driven mindset toward emerging engineering technologies, with a proven ability to integrate AI-augmented development into daily SRE workflows from scripting, debugging, and monitoring to incident retrospectives
- Actively leverage AI assistants and automation frameworks to synthesize complex technical data, streamline documentation, and automate repetitive tasks, consistently moving faster to solve reliability challenges and elevate the quality of our engineering output
- 2+ years deploying public cloud infrastructure (AWS, GCP, Azure) preferred including administering managed AWS services (EKS, OpenSearch, MSK, Batch, etc.)
- Experience with orchestration tooling such as Kubernetes
- Experience with Docker or similar container solution
- Bachelor's Degree or Master's degree in a technical field such as Computer Science, Information Technology Engineering or equivalent work experience
- Experience with the Agile software development methodology and collaboration with internal teams to deliver software and configuration artifacts
- Experience in bash scripting in addition to experience in higher-level scripting languages like Python or Node.js
- Experience contributing to projects through to completion as part of a team
- Experience deploying and managing distributed, microservice oriented applications at scale
- Be an enthusiastic learner, user, and advocate of our technologies
- Has desire to win as a team – make big things happen by working together and being open and willing to try new ideas
- Strong interpersonal and communications skills (written, verbal, & virtual) with ability to work in a team-oriented, collaborative environment
- Must have high degree of personal integrity and ability to maintain strict confidentiality
- Must have a strong drive, be self-motivated, logical, and have a keen attention to detail
And Ideally:
- Experience with platform development in compliance with FedRAMP requirements and related government standards, policies, regulations, etc.
- Experience deploying distributed, microservice oriented applications
- Experience with distributed monitoring tools such as: Datadog, Splunk, OTEL, etc.
- Experience with datastore technologies: Kafka, Elasticsearch, DynamoDB, RDS Aurora PostgreSQL
- Experience with Helm
- Experience with Go, Java/Kotlin, and/or Groovy
- Experience with Java build tools including Gradle
- 1+ years of operational experience with industry-leading "big data" services technologies
How would you rate this job post?
See what other professionals think about this role.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.