Lead DevOps Engineer
Job Details
About This Role
About the job
Experience: 6.00 + years
Salary: INR 4000000-5000000 / year (based on experience)
Expected Notice Period: 30 Days
Shift: (GMT+05:30) Asia/Kolkata (IST)
Opportunity Type: Office ()
Placement Type: Full Time Permanent position(Payroll and Compliance to be managed by: Skit.ai)
(*Note: This is a requirement for one of Uplers' client - Skit.ai)
What do you need for this opportunity?
Must have skills required:
AWS SA Professional, Azure Solutions, FinOps, GCP Professional Cloud Architect, KEDA, Livekit, MLOps, Pipecat, Twilio, AWS/GCP/Azure, CI/CD, DNS, IAM/RBAC, NAT, SIP/PSTN, Terraform, VPC/VNet design, WebRTC, Kubernetes
Skit.ai is Looking for:
Job Title: Lead DevOps Engineer
Location: Bengaluru (WFO)
Job Type: Full-time
The problem
Skit.ai runs autonomous voice agents for regulated enterprises — India''s largest banks and telcos, and US collections operations. Every call is a live distributed system: PSTN/SIP → media server → ASR → LLM → TTS → back, spread across three clouds and multiple vendors, with a conversational response budget measured in hundreds of milliseconds. The platform peaks at roughly 1 million calls per hour. Billed minutes grew 5,000x+ in eight months. At this scale, infrastructure is not a support function — latency, cost-per-minute, and auditability are product features. When infra degrades, a customer mid-sentence hears silence. We''re hiring a Lead DevOps Engineer to own this substrate and keep it ahead of the growth curve.
What you''ll own:
- Multicloud substrate: Production infrastructure across AWS, GCP, and Azure. Private interconnects (Direct Connect, Cloud Interconnect, ExpressRoute), transit/hub-spoke topologies, and cross-cloud latency managed as an explicit budget — p95 per hop in tens of milliseconds, not "best effort."
- Realtime media plane: Selfhosted LiveKit and SIP infrastructure at scale. Media servers are stateful you''ll design session-affine, event-driven autoscaling (KEDA-class) that survives traffic tripling within an hour.
- Modelserving infrastructure: GPU fleets (A100/H100/B200class) for selfhosted ASR and open-weight LLMs — inference optimization, prefix caching, sticky-session routing, sub-500ms TTFT budgets — alongside managed APIs (Vertex AI/Gemini, Bedrock, Azure). Vendor failover is your design, not your incident.
- Reliability & observability: OTelnative tracing (Grafana/Tempo stack), perturn latency attribution across telephony/ASR/LLM/TTS, automated incident response and self-healing. You''ll act as incident commander for infrastructure and write the runbooks you''d want at 3 a.m.
Cost engineering: Cost-per-minute is an SLO here. We cut per-minute serving cost ~18x in six months through caching, rightsizing, autoscaling, and workload re-architecture — you''ll own the next 10x.
-
Security & compliance: Zero Trust across clouds: private endpoints/PrivateLink, IAM/RBAC, secrets management with rotation, WAF/DDoS protection. Operate controls for SOC 2 and ISO/IEC 27001 working command of ISO/IEC 42001:2023 (AI management systems) — hands-on preferred, rigorous theoretical grounding acceptable. You''ll face bank and telecom auditors directly, including data-residency requirements.
-
Technical leadership: Terraformfirst IaC standards, productionreadiness reviews, mentoring SREs. "Lead" means you raise the floor of the whole team. Problems on our plate right now
-
Scaling stateful, selfhosted media servers past current concurrency ceilings — HPA on CPU doesn''t cut it
-
Migrating LLM inference from managed APIs to selfhosted openweight models on GPUs without breaking TTFT budgets
-
ASR, LLM, and telephony living in different clouds: interconnect topology that keeps the packet path short and private
-
Multiregion DR that satisfies bank audits without doubling spend If these read as interesting rather than terrifying, keep reading.
Must-have
- 6+ years handson cloud infrastructure 3+ years operating multiple clouds simultaneously in production deep expertise in at least two of AWS/GCP/Azure
- Realtime audio/video systems in production — WebRTC, SIP/PSTN, or streaming media you''ve debugged jitter, not just read about it
- Networking depth: VPC/VNet design, load balancing, DNS, NAT private connectivity (Direct Connect / Cloud Interconnect / ExpressRoute, PrivateLink / Private Service Connect) transit gateways and cross-cloud mesh
- Kubernetes at scale (EKS/GKE/AKS), Helm, and scaling stateful workloads service mesh familiarity (Istio/Linkerd)\ Infrastructure as Code: Terraform (non-negotiable) across multi-account/multi-project estates drift is a bug
- AI/ML serving in production: GPU allocation and scheduling, inference servers or serverless GPU platforms (vLLM / Triton / Modal / Baseten-class), streaming protocols (WebRTC, WebSocket, gRPC)
- Production STT/TTS/LLM API operations: streaming integrations, quota management, multi-vendor failover (Deepgram / Google / Azure / Whisper-class ASR ElevenLabs / Azure-class TTS)
- Security fundamentals: IAM/RBAC, secrets management (Vault or cloudnative), encryption and key rotation
- CI/CD: GitHub Actions or GitLab CI with security scanning integrated into the pipeline
Strong signal (nice-to-have)
- LiveKit, pipecat, Twilio, or comparable realtime platforms SIP trunking and PSTN integration
- KEDA or other eventdriven autoscaling used in anger
- MLOps: model versioning, canary and A/B rollout
- FinOps discipline: reserved/spot strategy, uniteconomics reporting
- Certifications: AWS SA Professional, GCP Professional Cloud Architect, Azure Solutions
Architect Expert
ISO/IEC 42001:2023 exposure
What We''re NOT Looking For
- Singlecloud depth with documentationlevel knowledge of the other two
- Toolchecklist DevOps without production AI/ML serving scars
- "Can learn quickly" as the primary qualification — this role needs day-one production credibility
- Anyone who has never traced a packet across a cloud boundary
How to apply for this opportunity?
Step 1: Click On Apply! And Register or Login on our portal. Step 2: Complete the Screening Form & Upload updated Resume Step 3: Increase your chances to get shortlisted & meet the client for the Interview!
About Uplers:
Our goal is to make hiring reliable, simple, and fast. Our role will be to help all our talents find and apply for relevant contractual onsite opportunities and progress in their career. We will support any grievances or challenges you may face during the engagement.
(Note: There are many more opportunities apart from this on the portal. Depending on the assessments you clear, you can apply for them as well).
So, if you are ready for a new challenge, a great work environment, and an opportunity to take your career to the next level, don't hesitate to apply today. We are waiting for you!