---
title: "Network Consultant II"
company: "Gruve"
company_url: "https://www.remjobs.works/companies/gruve"
url: "https://www.remjobs.works/job/gruve-network-consultant-ii-7237b91e-f9e7-4cc9-a177-74c0d1bfc01f"
apply_url: "https://gruve.ai/careers/?gh_jid=5407537008"
workplace: onsite
location: "Pune, Maharashtra, India"
employment_type: unspecified
seniority: mid
role: other
region: india
skills: ["gcp", "kubernetes", "linux", "llm", "python"]
date_posted: 2026-08-31T05:18:13.000Z
first_seen_by_remjobs: 2026-09-19T22:06:16.216Z
---

# Network Consultant II

**Gruve** · Pune, Maharashtra, India

Apply: https://gruve.ai/careers/?gh_jid=5407537008

## About the role

**About Gruve**

Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.

**Position summary: **

L2 escalation owner for the network/DMT track and shift anchor for the NOC pod, and L2 for the PulseAI infrastructure layer. Deep RCA on the Juniper fabric and edge network devices and Cisco firewalls, non-standard change execution and vendor TAC ownership — taking handoff from the L1 NOC bench and closing out or escalating to the Network Operations Consultant (L3) — plus remediation of GPU-server, node, fabric and OpenShift cluster-networking faults within Gruve's scope and execution of firmware/driver, switch-configuration and node-lifecycle changes within maintenance windows. Security investigations and PulseAI platform-layer remediation belong to the SOC pod and are not part of this role.

**Key responsibilities:**

- Act as shift anchor: own P1/P2 first response, open and run technical bridges until L3 engagement; own escalations from the NOC/DMT queue through resolution or structured L3 handoff.

- Complex network RCA (EVPN/BGP, fabric and edge health, interface/optics issues) across the in-scope Juniper data-center devices (QFX, EX, MX) and Cisco cdFMC/FMC, FTD/SRX; correlate fabric events with OpenShift node and cluster-network symptoms to separate network-side from cluster-side faults.

- Execute non-standard and complex changes beyond L1 authority under change control — fabric node additions, routing changes, firmware upgrades, firewall policy pushes — with senior oversight.

- Remediate PulseAI infrastructure incidents within Gruve's scope: GPU-server firmware and driver faults (within the change process), OpenShift node and cluster-networking issues (OVN-Kubernetes/CNI, node NotReady from hardware or network causes, MachineConfig drift), front-end and RoCEv2 back-end fabric connectivity, switch configuration within the PulseAI fabric; diagnose storage capacity/health and node-hardware faults and escalate to the vendor with complete diagnostics.

- Execute firmware and GPU driver updates, switch configuration changes and OpenShift node lifecycle operations (MachineConfig, node cordon/drain/reboot, node-pool changes) within agreed maintenance windows with the required customer notice; verify post-change node and cluster-network health with oc/kubectl.

- Own vendor TAC cases end to end — Juniper for fabric and edge; Cisco, and Cloudflare where a network issue touches that platform; OEM, neocloud-provider and storage-vendor cases for hardware — coordinate RMA/smart-hands, deliver daily status on Premium, and manage restoration-clock pauses correctly.

- Drive preventive maintenance: firmware/BIOS/driver currency, EOL/EOS tracking, spare posture — for network devices and PulseAI GPU and control-plane nodes.

- Own infrastructure observability for the NOC pod: GPU-cluster, fabric and OpenShift node health dashboards, log queries and alert thresholds in Grafana; contribute to the environment validation checklist at onboarding (telemetry reachability per switch, storage throughput, access path).

**Mandatory Qualifications:**

- 4–6 years NOC/network operations experience in data-center environments, including hands-on Junos (EVPN-VXLAN fabric) and/or Cisco cdFMC/FMC, FTD.

- Demonstrated independent RCA ownership on P1/P2/P3 network incidents; incident bridge experience; comfortable running a shift independently.

- Strong routing/switching troubleshooting (BGP, EVPN-VXLAN) and next-generation firewall operations; disciplined change-management practice.

- Hands-on Red Hat OpenShift / Kubernetes operations in production — node lifecycle and MachineConfig, operators, cluster networking (OVN-Kubernetes/CNI), storage, oc/kubectl troubleshooting — plus Linux (RHEL) administration.

- Working understanding of Kubernetes networking — CNI (Cilium), Services/Ingress/LoadBalancer and east-west flows — to troubleshoot fabric-to-cluster (GKE and OpenShift) connectivity end to end.

- Working understanding of GPU-server operations: NVIDIA GPU Operator and driver stack, DCGM-class telemetry, firmware/driver update procedures, RDMA/RoCEv2 NIC health and common GPU failure modes.

- Multi-vendor switch monitoring and troubleshooting (SNMP, syslog, streaming telemetry); out-of-band management proficiency.

**Preferred Qualifications:**

- JNCIP/JNCIS-DC or CCNP, or other professional-level data-center networking certification.

- Apstra or other intent-based networking exposure; IaC exposure; Python scripting / Ansible for operations automation.

- GKE cluster networking exposure (VPC-native, LoadBalancer/Gateway) and NetworkPolicy troubleshooting; working knowledge of GCP.

- Red Hat OpenShift Administration certification (EX280) or RHCSA/RHCE; CKA; exposure to AI/ML workload scheduling and GPU node pools on OpenShift.

- GPU-cluster performance troubleshooting — RoCEv2/PFC/ECN tuning, ECMP polarisation, NCCL-visible latency/jitter — and hands-on with high-performance fabric telemetry, NVLink/NVSwitch topologies and GPU-node network profiling.

- AI/HPC, neocloud or hyperscale data-center fabric exposure.

**Why Gruve**

At Gruve, we foster a culture of innovation, collaboration, and continuous learning. We are committed to building a diverse and inclusive workplace where everyone can thrive and contribute their best work. If you’re passionate about technology and eager to make an impact, we’d love to hear from you.

Gruve is an equal opportunity employer. We welcome applicants from all backgrounds and thank all who apply; however, only those selected for an interview will be contacted.

---

Source: Gruve's own career page, read by RemJobs. Canonical HTML version: https://www.remjobs.works/job/gruve-network-consultant-ii-7237b91e-f9e7-4cc9-a177-74c0d1bfc01f
