Site Reliability Engineer

Belgrade
On-site
Full-time

As a Site Reliability Engineer (SRE) at Inceptive, you will be joining a highly performant R&D team with a focus on testing and maintaining lab and testing environments. This role involves working on provisioning and optimizing the infrastructure for our services and getting the infrastructure-as-code (IaC) ready for pre-production testing and future production deployment. You will be working closely with application and system developers to ensure that our product is meeting the highest standards of reliability and performance. The work done now will serve as the foundation for our production environment, and will need to be designed to scale within a larger operations team.

Responsibilities

  • Infrastructure Monitoring
    • Develop and maintain hardware infrastructure monitoring system with alerting.
    • Develop and maintain database monitoring systems based on OTEL.
    • Develop and maintain application metrics collection and monitoring systems.
  • Deployment and Automation
    • Maintenance and provisioning of lab servers and the testing environment.
    • Configure and manage the network in the office/lab environments.
    • Maintain the CI/CD infrastructure.
  • Performance Optimization
    • Kernel parameter tuning for eliminating latency overhead and jitter in applications.
    • Conduct load testing and capacity planning.
  • Infrastructure Provisioning (IaC)
    • Develop and maintain infrastructure-as-code (IaC) solutions for reproducible provisioning of the production environment.
    • Use tools like Ansible and Helm to provision the infrastructure.
    • Provision and support of the computing infrastructure and settings.

Requirements

  • Professional Experience
    • Proven track record in SRE, DevOps, or system admin roles.
    • Experience managing hardware setups and coordinating with data center providers.
    • Strong analytical and problem-solving abilities.
    • Excellent communication and collaboration skills.
  • Technical Skills
    • Strong experience with Linux systems administration.
    • Knowledge of networking fundamentals.
    • Hands-on experience with automation and scripting tools, such as Ansible.
    • Knowledge of Docker and containerization.
    • Experience with orchestration tools such as Kubernetes or Hashicorp Nomad.
    • Experience with monitoring tools like Grafana, Prometheus, and Loki.
  • Bonus skills
    • Database administration.
    • Hardware and OS tuning knowledge.
    • Experience with low-latency system optimization and performance tuning.