Data & Knowledge Engineer
AI in this role
Job Description & Summary
The opportunity
Provide trusted, contextual and well-governed enterprise data and knowledge services that ground agentic workflows and improve their reliability.
What you will be doing
· Design and build ingestion, transformation and serving pipelines for structured and unstructured data.
· Create retrieval indexes, metadata models, semantic layers, knowledge graphs or data products as appropriate.
· Implement chunking, enrichment, lineage, quality and access-control patterns.
· Optimize retrieval quality, freshness, latency and cost with the AI engineering team.
· Integrate cloud and on-premises data sources for hybrid solutions.
· Support evaluation datasets, monitoring data and traceability requirements.
What we need from you
· 4+ years in data engineering, analytics engineering, information retrieval or knowledge platforms.
· Strong SQL and Python skills and experience with data pipelines, APIs and data modeling.
· Practical knowledge of vector search, embeddings, metadata, document processing and retrieval evaluation.
· Experience with enterprise security, data quality and hybrid data integration.
Relevant AI technologies and tooling
· Strong SQL and Python capability with practical experience in Spark and data engineering platforms such as Microsoft Fabric, Azure Data Factory, Databricks, Snowflake or equivalent.
· Hands-on experience processing structured and unstructured content, including parsing, OCR, chunking, enrichment, metadata extraction, lineage and incremental indexing.
· Experience with vector and hybrid search technologies such as Azure AI Search, PostgreSQL with pgvector, Elasticsearch, Pinecone, Weaviate, Milvus or equivalent.
· Understanding of embedding selection, semantic and lexical retrieval, metadata filtering, reranking, query transformation, evaluation datasets and retrieval quality metrics.
· Experience with graph and knowledge technologies such as Neo4j, RDF or property graphs, ontologies, entity resolution and GraphRAG patterns is desirable.
· Ability to implement secure hybrid data access, row or document-level permissions, data masking and traceable ingestion from cloud and on-premises repositories.
Measures of success
· Data freshness, quality and availability
· Retrieval relevance and traceability
· Speed of onboarding new knowledge sources
· Pipeline reliability and performance
· Compliance with data-access requirements
Key interfaces
· Other members of the AI Transformation & Agentic Systems Practice
· PwC sector, functional, cloud, cyber, risk, Responsible AI and change specialists
· Client business owners, product owners, technology teams and operational users
· Technology alliance and implementation partners where relevant
Contribution to the practice
· Support proposals, client workshops and market development appropriate to seniority.
· Contribute reusable methods, patterns, code, assets and lessons learned.
· Coach colleagues and participate in the capability’s continuous learning agenda.
· Uphold PwC quality, independence, confidentiality and risk-management requirements.
#LI-BS1 #LI-Hybrid
How we score this
Data & Knowledge Engineer at PwC scores 86 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.
AI Level 4. Building AI systems is the job itself: without AI, the role would not exist.
- AI Level 480 to 100
- AI Level 360 to 79
- AI Level 240 to 59
- AI Level 10 to 39
Bands come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How do you decide when an AI agent can act on its own versus asking for approval first?
- How do you think about the risk of an AI system in this kind of role failing silently?
- What are the limits of Pinecone that you've run into, and how did you work around them?
- What's a project where you used Weaviate hands-on?
- Walk me through how you've used pgvector in your day-to-day work.
Adapt your resume
- List these exact terms on your resume: AI Agents, AI Safety, Pinecone, Weaviate, and pgvector. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Get new AI jobs at AI Level 4+ by email
One email a week with the new AI jobs at AI Level 4+, each rated AI Level 1 to 4 for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Software Engineering roles rated AI Level 4 at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step