DevOps Engineer

Building infrastructure that heals itself.

I'm Chetan Kesare, a DevOps Engineer with 1.3+ years of experience who builds and operates cloud infrastructure end-to-end — from CI/CD pipeline design and container orchestration with Kubernetes and Helm, to Infrastructure as Code with Terraform and Ansible, to the observability and logging pipelines that keep production reliable. I'm comfortable across AWS and Oracle Cloud, and I care about high-availability, zero-downtime releases, fast incident response, and cutting MTTR — the kind of systems that stay up and stay boring.

View my work LinkedIn GitHub
Chetan Kesare

Experience

Resume

Download résumé (PDF)
cloud infrastructure / operations
CloudDrove · DevOps
Aug 2025 — Now
production
AWSECS / ALB
IAMleast privilege
OKEKubernetes
CIJenkins
30+microservices
4 envsDEV → PROD
24/7service uptime
MTTR ↓incident response
RUNTIME SURFACEHEALTH CHECK
AWS ECS / ALBstable
OCI OKE / MLOpshealthy
Redis / Oracle VM24/7
EC2 / RDS / ElastiCachepatched
deployment windowZERO-DOWNTIME
ecs service stable
iam policy secured
oke cluster healthy
deployment zero-downtime
Aug 2025 — Now

System Engineer — DevOps

CloudDrove Inc.

Project: Client Project (Ad-Tech AI Platform) · Toronto, Canada (Remote)

  • Provisioned and maintained AWS ECS services end-to-end — from architecting new services for feature rollouts to fixing existing services and Load Balancer (ALB) configurations — ensuring stable, production-ready deployments.
  • Managed AWS IAM security posture: created and maintained IAM roles and permission policies for services and users, enforcing least-privilege access and resolving permission-related incidents across the environment.
  • Drove AWS cost optimization initiatives, identifying and eliminating inefficient resource usage across compute and storage to reduce cloud spend.
  • Performed routine server patching and maintenance on EC2, RDS, and ElastiCache instances, keeping infrastructure secure, up to date, and audit-ready.
  • Administered a Kubernetes cluster on Oracle Cloud Infrastructure (OCI) hosting MLOps workloads — triaging pod failures, debugging cluster-level issues, and maintaining service health and uptime.
  • Managed Redis caching infrastructure on Oracle Cloud VMs, ensuring 24/7 availability through proactive monitoring and rapid resolution of caching failures and downtime.
  • Acted as the DevOps point of contact for client-facing engineering support, assisting with real-time application fixes, improvements, and deployment issues to keep client releases on track.
  • Built and maintained Jenkins CI/CD pipelines for a 30+ microservice platform, containerizing services with Docker and managing image builds and pushes to Amazon ECR, with Helm-based deployments to AWS ECS/EKS and Oracle OKE enabling zero-downtime releases across DEV, QA, STAGING, and PROD.
real-time application / frontend
Bluestock Fintech
Oct — Dec 2024
live data
MARKET / LIVE● CONNECTED
NIFTY 50 24,863.40+0.82%
real-time datastream stable
UI responseoptimized
productionbugs resolved
Oct — Dec 2024

Software Development Intern

Bluestock Fintech

Pune, India (Remote)

  • Developed and deployed UI features for a real-time, data-intensive stock-market platform, ensuring performance, accuracy, and stability under concurrent user load.
  • Resolved production bugs and improved interface responsiveness, reducing the frontend error rate.

Projects

Engineering case study

gitops deployment flow
1GitHub push → GitHub Actions
2Test & build Docker images
3Push images → Docker Hub
4Update Kubernetes manifests
5Argo CD syncs cluster state
6 Prometheus → Alertmanager → Discord
Animated ZeroBroker CI/CD, GitOps, Kubernetes, monitoring and alerting workflow

Animated architecture & deployment workflow

Kubernetes GitHub Actions Argo CD Docker Prometheus Grafana MongoDB Atlas

ZeroBroker — GitOps-Driven Real-Estate Microservices Platform

Architected and deployed a containerized real-estate application as a multi-service Kubernetes workload, combining automated CI/CD, GitOps delivery, and production-style observability into a single end-to-end platform.

What it demonstrates: React frontend and Node.js/TypeScript services are containerized with Docker and deployed to Kubernetes. A GitHub push triggers GitHub Actions to test and build images, publish them to Docker Hub, and update Kubernetes manifests; Argo CD then detects the desired-state change and synchronizes the cluster automatically.

Reliability & observability: Prometheus monitors cluster and application health, while alert rules detect unavailable replicas, crash-looping pods, failing endpoints, and down probes. Alertmanager routes actionable alerts to Discord, and Grafana provides dashboards for application health, resource usage, deployment state, endpoint performance, and alerts — demonstrating a complete build → deploy → observe → alert workflow.

View on GitHub ↗ Live Demo ↗

AI in my work

AutoPilot — an autonomous Kubernetes self-healing platform

A personal project exploring what happens when you give a cluster the ability to notice its own problems and fix them before a human has to page in.

Python Kubernetes Claude AI Jenkins Prometheus
SELF-HEALING DETECTION & REMEDIATION

Built an autonomous, AI-powered Kubernetes self-healing platform using the Kubernetes Python SDK, monitoring a containerized MERN microservices app (TaskFlow) on a self-hosted K3s cluster — auto-detecting 28+ failure types (CrashLoopBackOff, OOMKill, node pressure, ingress failures) and executing remediations, with Slack-gated approval for high-severity incidents.

PIPELINE FAILURE INTELLIGENCE

Engineered Jenkins CI/CD failure intelligence that classifies 24 pipeline failure patterns with auto-retry, pre-deploy health gates, and post-deploy verification via webhooks — plus a self-healing feedback loop with config drift detection and Prometheus metrics exposure.

DEPLOYMENT & INFRASTRUCTURE

Containerized and deployed the TaskFlow MERN application on an Oracle Cloud ARM64 K3s cluster, resolving NGINX reverse-proxy routing and multi-arch Docker build issues to achieve stable external production access.

View on GitHub ↗

Certifications

Learning

AWS Certified Cloud Practitioner

Amazon Web Services

In progress

Certified Kubernetes Administrator (CKA)

CNCF / The Linux Foundation

Planned

HashiCorp Certified: Terraform Associate

HashiCorp

Planned

Skills & tools

Skills

Cloud Platforms

AWS (EC2, ECS, EKS, ECR, RDS) AWS (VPC, IAM, ALB/ELB, Route 53) Auto Scaling Groups, CloudWatch Oracle Cloud Infrastructure (OKE)

Containers & Orchestration

Docker Kubernetes & Helm kubectl / K3s Microservices, Pod Lifecycle

CI/CD, IaC & Automation

Jenkins, GitHub Actions Terraform, Ansible Bash / Shell Scripting SonarQube, Nexus, Trivy

Observability & Monitoring

Prometheus, CloudWatch New Relic, Splunk Fluent Bit Incident Response, MTTR Reduction

Platforms & Practices

Linux, NGINX Load Balancing, High Availability Site Reliability Engineering (SRE) Disaster Recovery

Education

Education

2021 — 2024

B.E. Computer Engineering

Terna Engineering College, Navi Mumbai, India

CGPA: 8.23 / 10

2018 — 2021

Diploma in Computer Engineering

Nanasaheb Mahadik Polytechnic Institute, Sangli, India

Score: 92% — Top 5% of batch

2018

CBSE

Prakash Public School, Islampur, Maharashtra

Score: 70%

Balanced academics and athletics by winning two District and two Taluka football championships, competing at the Division level, and receiving recognition in the Hindi Olympiad.

Contact

Let's build reliable systems together.

Open to DevOps / SRE / Cloud Infrastructure roles where I can automate deployments, harden observability, and keep production reliable at scale.

Email chetankesare06@gmail.com ↗ LinkedIn chetan-kesare09 ↗ GitHub 0x1Luffy ↗ Phone +91 97307 28443 ↗

© 2026 Chetan Kesare — DevOps Engineer