Skip to main content
J
JobWave

Applied AI Engineer - Software Specialist Engineer II

Deloitte

Posted

  • Bengaluru, Karnataka, India

Job description

JOBWAVE VERIFIED

Applied AI Engineer II

Company: Deloitte
Requisition Code: 365121
Team: US Deloitte Technology Product Engineering
Type: Full Time
Experience: 5+ Years
GenAI / Agentic Experience: 3+ Years
Travel: ~10% (business/product dependent)
Immigration Sponsorship: Limited availability
Location: Not specified
Salary: Not specified

────────────────────

JOBWAVE INSIGHTS

Who is this for?

An experienced full-stack software engineer with 5+ years of engineering experience and 3+ years building production Generative AI and Agentic AI systems.

The role is designed for someone who can move beyond experimentation and build production-grade AI products, combining full-stack engineering, LLMs, RAG, multi-agent systems, cloud architecture, observability, evaluation, security, and FinOps.

Key Skills

  • Python
  • Node.js
  • C# / .NET
  • Java
  • React
  • Angular
  • SQL
  • NoSQL
  • REST APIs
  • GraphQL
  • Microservices
  • Object-Oriented Design
  • Data Structures & Algorithms
  • Generative AI
  • LLMs
  • Agentic AI
  • LangChain
  • LangGraph
  • RAG
  • Vector Databases
  • Multi-Agent Systems
  • Tool Calling
  • Human-in-the-Loop
  • LLM Evaluation
  • LangSmith
  • LangFuse
  • MLflow
  • PyTorch
  • TensorFlow
  • AWS
  • Azure
  • GCP
  • Azure OpenAI
  • AWS Bedrock
  • Vertex AI
  • Infrastructure as Code
  • FinOps
  • Token Optimization
  • CI/CD
  • DevSecOps
  • SRE
  • GitHub
  • Azure DevOps
  • SonarQube

Priority Skills

Highest priority:

  • Production Generative AI
  • Agentic AI / Multi-Agent Systems
  • Python or another major backend language
  • Full-stack development
  • RAG
  • Vector databases
  • LangChain / LangGraph
  • LLM evaluation
  • Cloud-native engineering
  • Microservices
  • System design
  • AI observability
  • Token and inference-cost optimization

Strong advantages:

  • LangSmith / LangFuse
  • MLflow
  • PyTorch / TensorFlow
  • Azure OpenAI
  • AWS Bedrock
  • Vertex AI
  • Infrastructure as Code
  • DevSecOps
  • SRE
  • Semantic caching
  • Hybrid search
  • Re-ranking
  • Quantized / smaller models
  • LLM output testing

What to Prepare

  • Full-stack system design
  • Object-Oriented Design
  • Data Structures & Algorithms
  • Microservice architecture
  • REST / GraphQL APIs
  • SQL and NoSQL databases
  • RAG architecture
  • Vector search
  • Hybrid search
  • Chunking strategies
  • Re-ranking
  • Multi-agent orchestration
  • LangGraph
  • LangChain
  • Tool calling
  • Agent state management
  • Human-in-the-loop workflows
  • LLM evaluation
  • AI observability
  • Token optimization
  • Inference latency optimization
  • Semantic caching
  • Model selection
  • Cloud AI platforms
  • AWS Bedrock
  • Azure OpenAI
  • Vertex AI
  • CI/CD
  • DevSecOps
  • SRE
  • Automated testing

What to Focus On

1. Production AI Engineering

The role is focused on production AI systems, not just prompt engineering.

Be prepared to explain the complete lifecycle:

User → Application → Agent → LLM → Tools / APIs → RAG → Data → Response → Evaluation → Monitoring

2. RAG Architecture

Understand the full RAG pipeline:

  • Document ingestion
  • Chunking
  • Embedding
  • Vector storage
  • Retrieval
  • Hybrid search
  • Re-ranking
  • Context construction
  • LLM generation
  • Evaluation

Be prepared to discuss how you reduce hallucinations and improve retrieval quality.

3. Multi-Agent Systems

Understand:

  • Agent orchestration
  • Agent roles
  • State management
  • Tool calling
  • Agent-to-agent communication
  • Failure handling
  • Human-in-the-loop escalation
  • Observability
  • Evaluation

Practice building a small multi-agent application using LangGraph or LangChain.

4. Full-Stack Engineering

Strong AI knowledge is not enough for this role.

Prepare for:

  • Backend API design
  • Microservices
  • Frontend integration
  • SQL / NoSQL databases
  • Vector databases
  • Authentication and authorization
  • API scalability
  • Error handling
  • System design

5. FinOps & Token Economics

Understand how AI applications create costs through:

  • Input tokens
  • Output tokens
  • Model selection
  • Number of LLM calls
  • Retrieval operations
  • Agent loops
  • Inference latency
  • Cloud infrastructure

Prepare optimization strategies such as:

  • Smaller models
  • Quantized models
  • Semantic caching
  • Prompt optimization
  • Context reduction
  • Request batching
  • Streaming
  • Model routing

6. Evals & Observability

Know how to measure AI system quality.

Prepare for:

  • LLM evaluation
  • Retrieval evaluation
  • Hallucination detection
  • Response quality
  • Agent traces
  • Latency monitoring
  • Token usage
  • Error monitoring

Understand tools such as:

  • LangSmith
  • LangFuse
  • MLflow

7. Cloud Architecture

Be prepared to compare:

  • AWS vs Azure vs GCP
  • AWS Bedrock
  • Azure OpenAI
  • Vertex AI

Understand how to deploy AI applications using cloud-native architecture.

8. DevSecOps & Engineering Quality

Prepare to discuss:

  • CI/CD
  • Automated testing
  • Security
  • Infrastructure as Code
  • Code quality
  • SRE
  • Monitoring
  • Vulnerability management
  • AI-specific testing

Potential Gaps

Candidates may need additional preparation if they have:

  • Strong software engineering but limited production GenAI experience
  • Prompt-engineering experience without production AI architecture
  • Limited multi-agent experience
  • Weak RAG knowledge
  • No vector database experience
  • Limited LLM evaluation experience
  • No AI observability experience
  • Weak cloud architecture knowledge
  • Limited FinOps/token optimization experience
  • Limited microservice architecture experience
  • Weak system design fundamentals
  • No DevSecOps experience
  • Limited automated LLM testing
  • Little experience with LangChain or LangGraph

Key experience requirement to verify: 5+ years of full-stack software engineering and 3+ years specifically focused on production GenAI/Agentic systems.

────────────────────

JOB DETAILS

Original Role Information

Role: Applied AI Engineer II
Requisition Code: 365121
Team: US Deloitte Technology Product Engineering
Travel: ~10% depending on business/product
Immigration Sponsorship: Limited availability

The role focuses on building production-grade full-stack products that integrate Generative AI, Agentic Workflows, RAG, multi-agent systems, and cloud-native technologies.

Responsibilities

Full-Stack & AI Architecture

  • Design full-stack web platforms.
  • Build backend microservices.
  • Integrate LLMs into production applications.
  • Build RAG pipelines.
  • Develop multi-agent systems.
  • Deploy production AI solutions.

FinOps & Optimization

  • Manage cloud infrastructure efficiency.
  • Optimize inference latency.
  • Reduce token consumption.
  • Monitor AI operational costs.
  • Balance technical performance with cost-to-value.
  • Optimize model and infrastructure selection.

Agentic SSDLC & DevSecOps

  • Implement CI/CD pipelines.
  • Apply spec-driven development.
  • Implement DevSecOps practices.
  • Use AI-assisted development tools.
  • Integrate automated testing into the development lifecycle.

Rapid Prototyping

  • Build lightweight prototypes.
  • Validate concepts quickly.
  • Gather early user feedback.
  • Iterate rapidly.
  • Avoid unnecessary upfront engineering when validating new concepts.

Cross-Functional Collaboration

  • Work with UX designers.
  • Collaborate with product managers.
  • Partner with business stakeholders.
  • Balance technical feasibility with user experience.
  • Align technical solutions with product strategy.

────────────────────

REQUIREMENTS

Experience

  • 5+ years of full-stack software engineering experience.
  • 3+ years building production Generative AI / Agentic AI systems.
  • 3+ years of cloud-native engineering experience.
  • Practical production experience rather than purely academic AI knowledge.

Education

Bachelor's degree in:

  • Computer Science
  • Software Engineering
  • Data Science
  • Machine Learning
  • Related quantitative field

Full-Stack Engineering

Experience with one or more:

  • Python
  • Node.js
  • C# / .NET
  • Java
  • React
  • Angular

Also required:

  • SQL databases
  • NoSQL databases
  • OOP / OOD
  • Data Structures & Algorithms
  • UML
  • Software architecture

Generative AI & Agentic Systems

  • Production LLM applications
  • OpenAI
  • Anthropic
  • Open-source LLMs
  • LangChain
  • LangGraph
  • PyTorch
  • TensorFlow
  • RAG
  • Vector databases
  • Multi-agent orchestration
  • Model evaluations
  • AI observability
  • LangSmith
  • LangFuse

Cloud & FinOps

  • AWS, Azure, or GCP
  • Azure OpenAI
  • AWS Bedrock
  • Vertex AI
  • Cloud-native architecture
  • Infrastructure as Code
  • Token optimization
  • Latency optimization
  • Cloud cost management

Engineering Tooling

  • GitHub
  • Azure DevOps
  • SonarQube
  • MLflow
  • CI/CD
  • DevSecOps
  • SRE
  • Lean / XP Agile

────────────────────

QUALIFICATIONS

  • Bachelor's degree in a relevant technical or quantitative discipline.
  • 5+ years full-stack software engineering experience.
  • 3+ years production GenAI/Agentic AI experience.
  • 3+ years cloud-native engineering experience.
  • Strong software architecture fundamentals.
  • Experience developing production LLM applications.
  • Experience with RAG and vector databases.
  • Experience with multi-agent systems.
  • Experience with AI evaluation and observability.
  • Understanding of FinOps and AI cost optimization.
  • Experience with CI/CD and DevSecOps.
  • Strong cross-functional collaboration skills.

────────────────────

INTERVIEW PREPARATION

Phase 1 — Full-Stack & System Design

Revise:

  • OOP/OOD
  • DSA
  • Clean architecture
  • Microservices
  • REST / GraphQL
  • SQL / NoSQL
  • ER diagrams
  • Sequence diagrams
  • Component diagrams

Practice designing an AI-powered application from frontend to database and LLM layer.

Phase 2 — RAG & Multi-Agent Systems

Build a small project using:

React → API → LangGraph/LangChain → RAG → Vector DB → LLM

Include:

  • Multiple agents
  • State management
  • Tool calling
  • Human approval
  • Error handling
  • Retrieval
  • Evaluation

Phase 3 — FinOps, Evals & Observability

Prepare to explain:

  • How you reduce token consumption
  • How you reduce LLM latency
  • How you select models based on cost and quality
  • How semantic caching works
  • How smaller/quantized models can reduce cost
  • How streaming improves perceived latency
  • How you evaluate LLM outputs
  • How you trace agent behavior

Practice with LangSmith, LangFuse, or MLflow.

Phase 4 — Architecture & Case Study

Prepare to defend trade-offs such as:

  • AWS Bedrock vs Azure OpenAI
  • Proprietary vs open-source models
  • RAG vs fine-tuning
  • Single-agent vs multi-agent
  • Synchronous vs asynchronous processing
  • Large vs smaller models
  • SQL vs NoSQL
  • Vector DB choices
  • Serverless vs containerized deployment

For each decision, explain:

Requirement → Options → Trade-offs → Decision → Cost → Performance → Security → Business Impact

How to apply

You will be taken to the employer’s site to complete your application.

Apply for this job (opens in a new tab)

Level up your career

Check out our career tips for practical advice on resumes, interviews, and job searching.

Browse Career Tips