GraphRAG Engineer
AI in this role
Engineer enterprise-grade GraphRAG, vector storage, and agentic AI pipelines to power developer platforms.
Business Area:
ITSeniority Level:
Mid-Senior levelJob Description:
At Cloudera, we empower people to transform complex data into clear and actionable insights. With as much data under management as the hyperscalers, we're the preferred data partner for the top companies in almost every industry. Powered by the relentless innovation of the open source community, Cloudera advances digital transformation for the world’s largest enterprises.
About the Team & Role
We are engineering an enterprise-grade Everything-as-Code (EaC) AI-First Platform that transforms modern enterprise operations through automated delivery pipelines, machine-readable specifications, and agentic intelligence. As a GraphRAG Engineer, you will own the semantic, vector, and graph storage layer powering the core context engine for our enterprise AI utilities and Internal Developer Portal.
Operating at the intersection of modern database administration, distributed event streaming, and generative AI pipelining, you will bridge our AWS MSK event mesh with downstream knowledge graphs and vector engines across AWS and GCP. You will lead the deployment of our SDLC Context Graph and GraphRAG Engine, enabling automated Change Advisory Board (CAB) compliance, semantic code/schema lineage tracking, and enterprise LLM proxy integrations.
As a GraphRAG Engineer you will:
- Graph & Vector Database Infrastructure: Provision, tune, and maintain production-grade clusters for Graph databases (Neo4j using Cypher, APOC, and causal clustering) and Vector storage engines (pgvector on PostgreSQL / AWS/GCP managed storage). Engineer high-throughput index structures, cosine similarity vector indexes, and query optimizations for sub-second responses.
- SDLC Context Graph & Lineage Pipelines: Build automated ingestion pipelines to parse Git repositories, ASTs, Jira issue links, Apache Avro schemas, and CI/CD metadata into a unified enterprise knowledge graph.
- GraphRAG Orchestration & Agentic Search: Connect distributed pipeline engines to hydrate hybrid retrievers (combining structured SQL, Cypher graph traversals, and dense vector embeddings) for AI-driven developer workflows and autonomous coding agents.
- Cyclic Agent Safeguards & Governance: Configure circuit breakers, confidence scoring thresholds, and step-limit constraints to restrict autonomous cyclic agent execution, protect token budgets, and prevent runaway execution loops.
- Prompts-as-Code & Enterprise LLM Gateway Integration: Integrate microservices and knowledge stores with the central Enterprise AI Gateway, maintaining version-controlled system prompt structures inside localized .ai/ spoke directories while adhering to DLP PII scrubbing rules and token rate limits.
- High Availability & FinOps: Implement automated failover, backup restoration, and multi-cloud storage tier cost controls across AWS and GCP environments.
We are excited if you have (Required Technical Expertise):
- Graph Databases: Deep operational and development experience with Neo4j (Cypher, APOC, causal clustering) or enterprise Knowledge Graphs.
- Vector Search & RAG: Proven expertise with pgvector (PostgreSQL), embeddings management, hybrid search techniques, and framework integrations (LangChain, LlamaIndex, or custom RAG pipelines).
- Database Administration & Cloud Storage: Hands-on experience managing relational (PostgreSQL) and graph databases across AWS and GCP cloud environments.
- Data Pipelining & Streaming: Proficiency in consuming Apache Avro payloads, streaming Kafka events (AWS MSK), and parsing structured/unstructured code and JSON artifacts.
- Agentic AI & Prompt Engineering: Practical understanding of Prompts-as-Code patterns, few-shot prompt optimization, and agent tool specification.
You may also have:
- Experience with Infrastructure-as-Code (Terraform) primitives, Kubernetes (EKS/GKE), Docker, and pull-based GitOps workflows.
- Exposure to HashiCorp Vault Transit encryption, OIDC keyless authentication, and zero-trust workload identities.
- Familiarity with OpenTelemetry (OTel) instrumentation for tracking vector search query latencies and LLM inference performance in Datadog or Grafana.
What you can expect from us:
Generous PTO Policy
Support work life balance with Unplugged Days
Flexible WFH Policy
Mental & Physical Wellness programs
Phone and Internet Reimbursement program
Access to Continued Career Development
Comprehensive Benefits and Competitive Packages
Employee Resource Groups
EEO/VEVRAA
#LI-LO1
#LI-HYBRID
How we rate this
GraphRAG Engineer at Cloudera rates 90 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.
Builds AI. The job is building AI systems.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How do you structure and test a prompt to get consistent output from a language model?
- How would you design a retrieval step so the model answers from real data instead of guessing?
- Tell me about a project where graph rag was part of your work. What did you do?
- Tell me about a project where vector databases was part of your work. What did you do?
- Tell me about a project where llm gateways was part of your work. What did you do?
Adapt your resume
- List these exact terms on your resume: Prompt Engineering, RAG, Graph RAG, Vector Databases, and LLM Gateways. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want an expert to read your CV for this job?
Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.
Get a free CV reviewGet new remote AI jobs (Builds AI ●●●●) by email
One email a week with the new remote AI jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Software Engineering roles that build AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step