Are you a Site Reliability Engineer eager to make critical AI infrastructure dependable, recoverable, and easy to operate?
Join Camplight, where your expertise will help engineering teams build and run trustworthy AI products at scale.
What you’ll be working on?
We are building an AI Observability platform for teams developing LLM-powered applications and agents. The platform gives engineers clear visibility into prompts, traces, tool calls, retrieval, performance, quality, and evaluation signals in production.
The platform is based on self-hosted Langfuse running on Kubernetes in AWS. Its architecture includes ClickHouse as the analytical store, PostgreSQL for metadata, Redis for ingestion queues, and Argo CD for delivery.
This platform will help product teams understand how their AI systems behave in the real world and enable them to build safer, more reliable, and more measurable AI experiences.
Your Role
Your role will involve taking ownership of the operational reliability of the AI Observability platform. You will build the foundations that make the service dependable: monitoring, alerting, service-level objectives, incident response, upgrades, backups, restore testing, and operational documentation.
You will diagnose and resolve failures across Kubernetes, ClickHouse, PostgreSQL, Redis, ingestion workers, and the delivery layer. You will work closely with the Platform Engineer, who owns platform capability and configuration, while you own service health, recoverability, safe migrations, and operational readiness.
About Camplight
We build self-organizing technical teams, offer software development services, and work with businesses and entrepreneurs to create new products.
With over 300 successful software projects, some ongoing for over 8 years, we strive for long-term success for our partners.
By following the principles of self-management and organizing as a cooperative, we achieve 95% satisfaction among them.
We seek the best talent to join us and value transparency, collaboration, trust, responsibility, and innovation.
When joining Camplight, you can become a co-owner of the cooperative, allowing you to steer the business and share in the rewards of our collective success.
What are we looking for?
- Ownership mindset: We want individuals who care deeply about service quality. You take responsibility for root causes, improve the system, and leave clear paths for others to operate it.
- Technical expertise: You know how to operate distributed systems in production, design useful monitoring, and plan for failure before it happens.
- Communication skills: You can write clear runbooks, explain operational risks, lead calm incident response, and collaborate effectively across engineering teams.
Requirements
- 4+ years of hands-on SRE, Platform Engineering, DevOps, or Infrastructure Engineering experience.
- Strong production experience with Kubernetes on AWS, ideally EKS.
- Experience operating self-hosted software through Helm, Kubernetes manifests, Argo CD, or comparable GitOps tooling.
- Production experience with ClickHouse, including replication, Keeper quorum, shard and replica topology, backups, and restores.
- Experience operating PostgreSQL and Redis in production.
- Experience diagnosing queue, backpressure, worker-drain, and ingestion issues.
- Experience building monitoring, dashboards, alerting, and meaningful service-level objectives.
- Experience planning and executing safe upgrades, schema migrations, rollback plans, and restore rehearsals.
- Strong infrastructure-as-code and automation habits.
- Experience creating runbooks, SOPs, post-incident reviews, and handover documentation.
- Strong written and spoken English.
What do we offer?
We focus on health, wealth, and empowering relationships:
- Fully remote work with flexible work hours
- Competitive salary
- Opportunity to become a co-owner of the cooperative
- Individual career development plan
- Friendly team and company culture
- Prioritization of mental and physical health in the workplace, with the freedom to make decisions about oneself, supported by peers committed to a healthy lifestyle.
- Empowering relationships for engineering alongside colleagues who cherish growth mindsets in a unique environment that blends service and product craftsmanship.
What does the interview process look like?
- Initial Interview: We’ll start with a friendly 45-minute cultural and technical interview. Two members of our team will assess your cultural fit, past experience, engineering expertise, the major challenges you’ve tackled, and discuss your ideal workspace.
- You can choose between two Technical Deep Dive options:
- Homework Assignment: If there’s a match, we’ll provide a brief homework assignment designed to take around 2 hours to complete. This will be followed by a 1-hour technical interview to discuss the homework and conduct a technical deep dive.
- Pair Programming: If you prefer not to do a homework assignment, we’ll have a 2-hour technical deep dive session focused on a practical Site Reliability Engineering scenario.
Join us at Techlight to connect with like-minded builders and grow alongside a trusted network of experts. Techlight is an invite-only community for developers and industry experts to exchange ideas, share experiences, and learn from one another.
Regardless of the outcome, we will provide you with constructive feedback to help you grow.
PEA Registration Number and Date of the Certificate 4003 / 10.10.2025