L1
Operations Engineer
Location: Bangalore
Experience
Required: 0 – 2 Years
Employment
Type: Full-Time | Shift-Based (24x7 Rotational)
About Zybisys
Zybisys is a technology company
that helps banks, financial institutions, and FinTech businesses build and run
secure, reliable, and high-performance technology platforms. We work closely
with some of India's leading stock brokers to manage their cloud
infrastructure, cybersecurity, platform operations, and observability. With
deep expertise in the Capital Markets domain, we focus on simplifying complex
technology, improving operational resilience, and helping our customers innovate
with confidence.
Job Description
The L1 Operations Engineer is
the first line of defence in Zybisys's 24x7 managed cloud operations. This role
is responsible for continuous monitoring of multi-region Azure cloud and hybrid
on-premises infrastructure, timely triage of alerts, first-line incident
response, and ensuring all issues are accurately logged, prioritised, and
escalated within SLA thresholds.
This is a shift-based
operational role requiring strong attention to detail, disciplined runbook
execution, and clear communication during incidents. The ideal candidate is
technically curious, process-oriented, and comfortable working across cloud
monitoring tools, ITSM platforms, and infrastructure dashboards in a fast-paced
managed services environment.
Key Responsibilities
Infrastructure Monitoring &
Alert Management
· Perform continuous 24x7
monitoring of Azure cloud and hybrid datacenter infrastructure using
Prometheus, Grafana, Azure Monitor, and related observability dashboards.
· Triage incoming alerts — assess
severity, validate against known patterns, and determine whether to resolve at
L1 or escalate to L2 within defined SLA thresholds.
· Execute approved runbooks and
SOPs for all known alert categories; document actions taken for every incident
with accurate timestamps and observations.
· Monitor health and availability
of compute (VMs), storage, network links, VPN tunnels, and platform services
across cloud and on-premises environments.
· Track sFlow and NetFlow
dashboards for network traffic anomalies; flag unusual patterns to the L2 team
for deeper investigation.
Incident Logging & ITSM
Management
· Log all incidents, service
requests, and alerts in the ITSM platform with complete and accurate details —
symptoms, affected components, priority, and initial actions taken.
· Update ticket status throughout
the incident lifecycle; ensure no incident is left without a current status
update beyond the defined response window.
· Coordinate with L2 engineers
during escalations — provide clear handover notes including timeline, alert
context, initial diagnostics, and business impact assessment.
· Follow the priority matrix
strictly: P1 (15 min), P2 (30 min), P3 (4 hr), P4 (8 hr) response and
escalation thresholds.
Platform & Service Health
Checks
· Execute scheduled shift health
checks across all managed platforms — Azure resources, on-premises servers,
network devices, security appliances, and employee services.
· Verify availability and
performance of core services: DNS, DHCP, NTP, Active Directory, and M365
platform components.
· Monitor security platform
dashboards (firewalls, EDR, proxy services) for health status and alert flags;
escalate anomalies per defined procedures.
· Review Azure Cost Management
dashboards for unusual consumption spikes and flag to the lead for review.
Routine Operations &
Maintenance Support
· Execute scheduled batch jobs,
backup verifications, replication checks, and housekeeping tasks as per the
operational calendar.
· Support L2 and Specialist
engineers during planned maintenance windows, patch cycles, and change
activities — providing monitoring coverage and rollback readiness.
· Validate post-change
infrastructure health after every approved change; raise a flag immediately if
anomalies are detected post-implementation.
Documentation & Knowledge
Management
· Maintain precise shift handover
reports — open tickets, ongoing incidents, recent changes, and watch-items for
the next shift.
· Contribute to the knowledge base
by documenting recurring alert patterns, resolution steps, and workarounds for
L1-resolvable issues.
· Flag gaps in runbooks or SOPs to
the operations lead so that documentation is continuously improved.
Required Skills & Experience
Cloud & Infrastructure
Fundamentals
· 2 – 5 years of experience in IT
operations, infrastructure support, or cloud managed services.
· Working knowledge of Microsoft
Azure: Azure Portal navigation, VM status checks, resource monitoring, and
basic troubleshooting using Azure Monitor and Log Analytics.
· Hands-on familiarity with
Windows Server (2016/2019/2022) and Linux (RHEL/Ubuntu) — service management,
log file reading, process monitoring, and basic fault diagnosis.
· Understanding of hybrid
infrastructure models — on-premises datacenter integrated with Azure cloud via
ExpressRoute or VPN.
· Basic familiarity with storage
concepts: disk performance thresholds, capacity monitoring, backup job status,
and replication health checks.
Monitoring & Observability
Tools
· Experience reading and
interpreting Prometheus metrics and Grafana dashboards — understanding panel
thresholds, alert states, and trend data.
· Familiarity with Azure Monitor
alerts and Log Analytics at a basic level — understanding alert rules, severity
levels, and affected resources.
· Ability to read sFlow or NetFlow
traffic dashboards for high-level network health assessment.
· APM dashboard familiarity —
understanding response time trends, error rates, and service health indicators
from tools such as Dynatrace, AppDynamics, or Azure Application Insights.
Networking Fundamentals
· Good understanding of core
networking concepts: TCP/IP, DNS, DHCP, NTP, VLANs, and basic routing.
· Practical diagnostic skills:
ping, traceroute, nslookup, netstat — to perform first-level connectivity
checks and provide meaningful diagnostics to L2.
· Familiarity with VPN tunnel
health monitoring and firewall status dashboards at an operational level.
ITSM & Process Discipline
· Experience working with ITSM
platforms (ServiceNow, Zoho Desk, Freshservice, or equivalent) for incident
logging, ticket updates, and escalation workflows.
· Understanding of ITIL incident
management concepts — priority, impact, urgency, escalation paths, and SLA
tracking.
· Ability to follow runbooks and
SOPs precisely and consistently, including under pressure during major incident
scenarios.
· Clear and concise written
communication for incident tickets, shift handover notes, and escalation
summaries.
Soft Skills & Work Style
· Comfortable working in a 24x7
rotational shift environment including night shifts, weekends, and public
holidays.
· High attention to detail —
accurate logging, precise documentation, and consistent process adherence.
· Team-oriented with a proactive
attitude toward learning and improving operational knowledge.
· Ability to stay calm and
systematic during high-pressure P1/P2 incident scenarios.
Certification (Preferred)
Domain
| Certification
|
Microsoft Azure
| AZ-900 (Azure
Fundamentals) | AZ-104 (Administrator — working towards)
|
Service Management
| ITIL v4 Foundation
|
Networking
| CompTIA Network+ | Cisco CCNA (advantageous)
|
Monitoring
| Grafana Certified Associate
(advantageous)
|