Offensive AI Engineer – Frontier AI Security & Testing, VP
AI in this role
Offensive AI Engineer – Frontier AI Security & Testing
Who we are looking for
State Street Global Cybersecurity (GCS) is looking for a hands-on Offensive AI Engineer to join Frontier AI Security & Testing (FAST). FAST is a new research and development team within Regulatory Assurance, Penetration Testing & Offensive Research. This role works at the intersection of frontier AI, cloud engineering, data engineering and offensive security. You will design, deploy and run the secure AWS-based platforms, evaluation harnesses and data pipelines that let State Street use frontier AI models safely as an offensive security capability.
This is an engineering role, not an advisory or architecture-only one. You will build, configure, deploy, troubleshoot and automate. You will work directly with frontier models, agentic frameworks and enterprise APIs. The work also includes putting technical controls in place so that AI is used in an auditable, governed way in a highly regulated financial services environment.
Why this role is important to us
State Street is expanding its use of frontier AI across cybersecurity to strengthen resilience, speed up adversary emulation and improve how defenses are validated. FAST decides which models can be trusted, shows with evidence what they can do, builds the platform they run on, and turns model capability into tools that red team, penetration testing, purple team and threat hunt teams can use. As an Offensive AI Engineer, you will help build the execution, evaluation and control layer that lets these capabilities run safely at enterprise scale.
What you will be responsible for
As an Offensive AI Engineer you will:
AWS AI infrastructure and deployment
- Design, engineer and deploy secure, isolated AWS environments to host and access frontier AI models. This includes Amazon Bedrock, SageMaker, private endpoint connectivity (PrivateLink/VPC endpoints), EKS/ECS, Lambda and API Gateway.
- Build and maintain infrastructure-as-code (Terraform or CloudFormation) and CI/CD pipelines for AI workloads, with repeatable and auditable deployment patterns.
- Engineer network segmentation, identity and access controls, secrets management and egress restrictions that create enforceable trust boundaries around agentic AI activity.
Frontier model engineering and evaluation
- Integrate with frontier model provider APIs and SDKs from major commercial vendors. Build model-agnostic abstraction layers that allow fast model onboarding and side-by-side comparison.
- Design and build standardized evaluation harnesses and benchmarking frameworks. These should measure model capability, precision, false-positive rate, reliability, cost and failure modes across offensive security tasks in a repeatable way.
- Develop agentic workflows, tool-use integrations and multi-agent orchestration patterns (e.g., MCP, LangGraph, Strands, or similar) that support reconnaissance, attack-path analysis and adversary simulation in controlled lab environments.
- Stress-test models against known AI failure modes, including prompt injection, jailbreak susceptibility, data leakage, tool misuse, scope drift and unsafe emergent behavior.
Data science and Databricks
- Design and build Databricks pipelines (PySpark/SQL) to ingest, curate and analyze model telemetry, evaluation results and offensive security data sets.
- Apply data science methods to evaluation results, including statistical comparison, scoring methodology, trend analysis and regression detection across model versions.
- Build data products and dashboards that give leadership evidence of model performance, cost and risk.
API integration
- Build secure, scalable integrations between AI platforms and enterprise systems, security tools and data sources using REST APIs, event-driven architectures and tool-execution brokers.
- Develop reusable API patterns and connectors for tool invocation, context injection and output validation.
AI controls, logging and governance
- Implement technical controls that limit agent autonomy where needed. These include guardrails, human-in-the-loop checkpoints, rate and scope limits, kill-switch mechanisms and least-privilege tool access.
- Engineer full logging, monitoring and observability that capture agent inputs, outputs, tool calls and decision paths to support auditability and post-incident analysis.
- Align AI use with enterprise Responsible AI, model risk management, legal and compliance requirements. Produce technical evidence suitable for audit and regulatory review.
- Document architectures, control implementations and evaluation methods to a standard that holds up under internal and external scrutiny.
What we value
These skills will help you succeed in this role:
- A builder's mindset: you are comfortable taking an idea from prototype to a hardened, production-grade capability.
- Deep hands-on experience deploying AI-enabled applications on AWS, particularly Amazon Bedrock, SageMaker, Lambda, API Gateway, EKS/ECS, IAM, KMS and VPC networking.
- Direct experience working with frontier LLMs through vendor APIs, including prompt and context engineering, tool use and function calling, agentic workflows and model evaluation.
- Strong Python skills, plus experience with GenAI and agentic frameworks (LangChain, LangGraph, Strands, CrewAI, Pydantic, MCP, or similar).
- Hands-on Databricks experience (PySpark, SQL, Delta Lake, workflows) and a working grounding in data science and statistical analysis.
- Experience designing and consuming REST APIs and event-driven integrations.
- A working understanding of AI security risks and controls, such as the OWASP Top 10 for LLM Applications, MITRE ATLAS and the NIST AI RMF.
- Familiarity with offensive security concepts, adversary tactics and frameworks such as MITRE ATT&CK (strongly preferred).
- Strong critical thinking and problem-solving skills, and the ability to explain complex technical results clearly to engineers, leadership and risk partners.
- The ability to research, learn and apply new models, tools and techniques quickly as the frontier AI landscape changes.
Education & Preferred Qualifications
- Bachelor's degree in Computer Science, Engineering, Data Science, Information Systems, Artificial Intelligence, or equivalent practical experience.
- 5+ years of hands-on experience in software, cloud, data or security engineering.
- 2+ years of hands-on experience building and deploying GenAI, LLM or agentic AI solutions on cloud platforms, preferably AWS.
- Experience with infrastructure-as-code (Terraform preferred) and CI/CD tooling.
- AWS certifications (e.g., Solutions Architect, Machine Learning Specialty, Security Specialty) and/or Databricks certifications are a plus.
- Offensive security certifications (e.g., OSCP, OSEP, GPEN) or penetration testing experience are a plus.
- Experience in financial services or another highly regulated industry is preferred.
Salary Range:
$120,000 - $202,500 AnnualThe range quoted above applies to the role in the primary location specified. If the candidate would ultimately work outside of the primary location above, the applicable range could differ.
Employees are eligible to participate in State Street’s comprehensive benefits program, which includes: our retirement savings plan (401K) with company match; insurance coverage including basic life, medical, dental, vision, long-term disability, and other optional additional coverages; paid-time off including vacation, sick leave, short term disability, and family care responsibilities; access to our Employee Assistance Program; incentive compensation including eligibility for annual performance-based awards (excluding certain sales roles subject to sales incentive plans); and, eligibility for certain tax advantaged savings plans.
For a full overview, visit https://hrportal.ehr.com/statestreet/Home.
About State Street
Across the globe, institutional investors rely on us to help them manage risk, respond to challenges, and drive performance and profitability. We keep our clients at the heart of everything we do, and smart, engaged employees are essential to our continued success.
We are committed to fostering an environment where every employee feels valued and empowered to reach their full potential. As an essential partner in our shared success, you’ll benefit from inclusive development opportunities, flexible work-life support, paid volunteer days, and vibrant employee networks that keep you connected to what matters most. Join us in shaping the future.
As an Equal Opportunity Employer, we consider all qualified applicants for all positions without regard to race, creed, color, religion, national origin, ancestry, ethnicity, age, disability, genetic information, sex, sexual orientation, gender identity or expression, citizenship, marital status, domestic partnership or civil union status, familial status, military and veteran status, and other characteristics protected by applicable law.
Discover more information on jobs at StateStreet.com/careers
Read our CEO Statement
Job Application Disclosure:
It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.
How we rate this
Offensive AI Engineer – Frontier AI Security & Testing, VP at State Street rates 89 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.
Builds AI. The job is building AI systems.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How do you decide when an AI agent can act on its own versus asking for approval first?
- How do you decide that one model's output is better than another's for a given task?
- How do you think about the risk of an AI system in this kind of role failing silently?
- What's a project where you used LangChain hands-on?
- Walk me through how you've used LangGraph in your day-to-day work.
Adapt your resume
- List these exact terms on your resume: AI Agents, AI Evaluation, AI Safety, LangChain, and LangGraph. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want an expert to read your CV for this job?
Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.
Get a free CV reviewGet new AI engineer jobs (Builds AI ●●●●) by email
One email a week with the new AI engineer jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Software Engineering roles that build AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step