Applied AI Engineer - Software Specialist Engineer II
Deloitte
Posted
- Bengaluru, Karnataka, India
Job description
JOBWAVE VERIFIED
Applied AI Engineer II
Company: Deloitte
Requisition Code: 365121
Team: US Deloitte Technology Product Engineering
Type: Full Time
Experience: 5+ Years
GenAI / Agentic Experience: 3+ Years
Travel: ~10% (business/product dependent)
Immigration Sponsorship: Limited availability
Location: Not specified
Salary: Not specified
────────────────────
JOBWAVE INSIGHTS
Who is this for?
An experienced full-stack software engineer with 5+ years of engineering experience and 3+ years building production Generative AI and Agentic AI systems.
The role is designed for someone who can move beyond experimentation and build production-grade AI products, combining full-stack engineering, LLMs, RAG, multi-agent systems, cloud architecture, observability, evaluation, security, and FinOps.
Key Skills
- Python
- Node.js
- C# / .NET
- Java
- React
- Angular
- SQL
- NoSQL
- REST APIs
- GraphQL
- Microservices
- Object-Oriented Design
- Data Structures & Algorithms
- Generative AI
- LLMs
- Agentic AI
- LangChain
- LangGraph
- RAG
- Vector Databases
- Multi-Agent Systems
- Tool Calling
- Human-in-the-Loop
- LLM Evaluation
- LangSmith
- LangFuse
- MLflow
- PyTorch
- TensorFlow
- AWS
- Azure
- GCP
- Azure OpenAI
- AWS Bedrock
- Vertex AI
- Infrastructure as Code
- FinOps
- Token Optimization
- CI/CD
- DevSecOps
- SRE
- GitHub
- Azure DevOps
- SonarQube
Priority Skills
Highest priority:
- Production Generative AI
- Agentic AI / Multi-Agent Systems
- Python or another major backend language
- Full-stack development
- RAG
- Vector databases
- LangChain / LangGraph
- LLM evaluation
- Cloud-native engineering
- Microservices
- System design
- AI observability
- Token and inference-cost optimization
Strong advantages:
- LangSmith / LangFuse
- MLflow
- PyTorch / TensorFlow
- Azure OpenAI
- AWS Bedrock
- Vertex AI
- Infrastructure as Code
- DevSecOps
- SRE
- Semantic caching
- Hybrid search
- Re-ranking
- Quantized / smaller models
- LLM output testing
What to Prepare
- Full-stack system design
- Object-Oriented Design
- Data Structures & Algorithms
- Microservice architecture
- REST / GraphQL APIs
- SQL and NoSQL databases
- RAG architecture
- Vector search
- Hybrid search
- Chunking strategies
- Re-ranking
- Multi-agent orchestration
- LangGraph
- LangChain
- Tool calling
- Agent state management
- Human-in-the-loop workflows
- LLM evaluation
- AI observability
- Token optimization
- Inference latency optimization
- Semantic caching
- Model selection
- Cloud AI platforms
- AWS Bedrock
- Azure OpenAI
- Vertex AI
- CI/CD
- DevSecOps
- SRE
- Automated testing
What to Focus On
1. Production AI Engineering
The role is focused on production AI systems, not just prompt engineering.
Be prepared to explain the complete lifecycle:
User → Application → Agent → LLM → Tools / APIs → RAG → Data → Response → Evaluation → Monitoring
2. RAG Architecture
Understand the full RAG pipeline:
- Document ingestion
- Chunking
- Embedding
- Vector storage
- Retrieval
- Hybrid search
- Re-ranking
- Context construction
- LLM generation
- Evaluation
Be prepared to discuss how you reduce hallucinations and improve retrieval quality.
3. Multi-Agent Systems
Understand:
- Agent orchestration
- Agent roles
- State management
- Tool calling
- Agent-to-agent communication
- Failure handling
- Human-in-the-loop escalation
- Observability
- Evaluation
Practice building a small multi-agent application using LangGraph or LangChain.
4. Full-Stack Engineering
Strong AI knowledge is not enough for this role.
Prepare for:
- Backend API design
- Microservices
- Frontend integration
- SQL / NoSQL databases
- Vector databases
- Authentication and authorization
- API scalability
- Error handling
- System design
5. FinOps & Token Economics
Understand how AI applications create costs through:
- Input tokens
- Output tokens
- Model selection
- Number of LLM calls
- Retrieval operations
- Agent loops
- Inference latency
- Cloud infrastructure
Prepare optimization strategies such as:
- Smaller models
- Quantized models
- Semantic caching
- Prompt optimization
- Context reduction
- Request batching
- Streaming
- Model routing
6. Evals & Observability
Know how to measure AI system quality.
Prepare for:
- LLM evaluation
- Retrieval evaluation
- Hallucination detection
- Response quality
- Agent traces
- Latency monitoring
- Token usage
- Error monitoring
Understand tools such as:
- LangSmith
- LangFuse
- MLflow
7. Cloud Architecture
Be prepared to compare:
- AWS vs Azure vs GCP
- AWS Bedrock
- Azure OpenAI
- Vertex AI
Understand how to deploy AI applications using cloud-native architecture.
8. DevSecOps & Engineering Quality
Prepare to discuss:
- CI/CD
- Automated testing
- Security
- Infrastructure as Code
- Code quality
- SRE
- Monitoring
- Vulnerability management
- AI-specific testing
Potential Gaps
Candidates may need additional preparation if they have:
- Strong software engineering but limited production GenAI experience
- Prompt-engineering experience without production AI architecture
- Limited multi-agent experience
- Weak RAG knowledge
- No vector database experience
- Limited LLM evaluation experience
- No AI observability experience
- Weak cloud architecture knowledge
- Limited FinOps/token optimization experience
- Limited microservice architecture experience
- Weak system design fundamentals
- No DevSecOps experience
- Limited automated LLM testing
- Little experience with LangChain or LangGraph
Key experience requirement to verify: 5+ years of full-stack software engineering and 3+ years specifically focused on production GenAI/Agentic systems.
────────────────────
JOB DETAILS
Original Role Information
Role: Applied AI Engineer II
Requisition Code: 365121
Team: US Deloitte Technology Product Engineering
Travel: ~10% depending on business/product
Immigration Sponsorship: Limited availability
The role focuses on building production-grade full-stack products that integrate Generative AI, Agentic Workflows, RAG, multi-agent systems, and cloud-native technologies.
Responsibilities
Full-Stack & AI Architecture
- Design full-stack web platforms.
- Build backend microservices.
- Integrate LLMs into production applications.
- Build RAG pipelines.
- Develop multi-agent systems.
- Deploy production AI solutions.
FinOps & Optimization
- Manage cloud infrastructure efficiency.
- Optimize inference latency.
- Reduce token consumption.
- Monitor AI operational costs.
- Balance technical performance with cost-to-value.
- Optimize model and infrastructure selection.
Agentic SSDLC & DevSecOps
- Implement CI/CD pipelines.
- Apply spec-driven development.
- Implement DevSecOps practices.
- Use AI-assisted development tools.
- Integrate automated testing into the development lifecycle.
Rapid Prototyping
- Build lightweight prototypes.
- Validate concepts quickly.
- Gather early user feedback.
- Iterate rapidly.
- Avoid unnecessary upfront engineering when validating new concepts.
Cross-Functional Collaboration
- Work with UX designers.
- Collaborate with product managers.
- Partner with business stakeholders.
- Balance technical feasibility with user experience.
- Align technical solutions with product strategy.
────────────────────
REQUIREMENTS
Experience
- 5+ years of full-stack software engineering experience.
- 3+ years building production Generative AI / Agentic AI systems.
- 3+ years of cloud-native engineering experience.
- Practical production experience rather than purely academic AI knowledge.
Education
Bachelor's degree in:
- Computer Science
- Software Engineering
- Data Science
- Machine Learning
- Related quantitative field
Full-Stack Engineering
Experience with one or more:
- Python
- Node.js
- C# / .NET
- Java
- React
- Angular
Also required:
- SQL databases
- NoSQL databases
- OOP / OOD
- Data Structures & Algorithms
- UML
- Software architecture
Generative AI & Agentic Systems
- Production LLM applications
- OpenAI
- Anthropic
- Open-source LLMs
- LangChain
- LangGraph
- PyTorch
- TensorFlow
- RAG
- Vector databases
- Multi-agent orchestration
- Model evaluations
- AI observability
- LangSmith
- LangFuse
Cloud & FinOps
- AWS, Azure, or GCP
- Azure OpenAI
- AWS Bedrock
- Vertex AI
- Cloud-native architecture
- Infrastructure as Code
- Token optimization
- Latency optimization
- Cloud cost management
Engineering Tooling
- GitHub
- Azure DevOps
- SonarQube
- MLflow
- CI/CD
- DevSecOps
- SRE
- Lean / XP Agile
────────────────────
QUALIFICATIONS
- Bachelor's degree in a relevant technical or quantitative discipline.
- 5+ years full-stack software engineering experience.
- 3+ years production GenAI/Agentic AI experience.
- 3+ years cloud-native engineering experience.
- Strong software architecture fundamentals.
- Experience developing production LLM applications.
- Experience with RAG and vector databases.
- Experience with multi-agent systems.
- Experience with AI evaluation and observability.
- Understanding of FinOps and AI cost optimization.
- Experience with CI/CD and DevSecOps.
- Strong cross-functional collaboration skills.
────────────────────
INTERVIEW PREPARATION
Phase 1 — Full-Stack & System Design
Revise:
- OOP/OOD
- DSA
- Clean architecture
- Microservices
- REST / GraphQL
- SQL / NoSQL
- ER diagrams
- Sequence diagrams
- Component diagrams
Practice designing an AI-powered application from frontend to database and LLM layer.
Phase 2 — RAG & Multi-Agent Systems
Build a small project using:
React → API → LangGraph/LangChain → RAG → Vector DB → LLM
Include:
- Multiple agents
- State management
- Tool calling
- Human approval
- Error handling
- Retrieval
- Evaluation
Phase 3 — FinOps, Evals & Observability
Prepare to explain:
- How you reduce token consumption
- How you reduce LLM latency
- How you select models based on cost and quality
- How semantic caching works
- How smaller/quantized models can reduce cost
- How streaming improves perceived latency
- How you evaluate LLM outputs
- How you trace agent behavior
Practice with LangSmith, LangFuse, or MLflow.
Phase 4 — Architecture & Case Study
Prepare to defend trade-offs such as:
- AWS Bedrock vs Azure OpenAI
- Proprietary vs open-source models
- RAG vs fine-tuning
- Single-agent vs multi-agent
- Synchronous vs asynchronous processing
- Large vs smaller models
- SQL vs NoSQL
- Vector DB choices
- Serverless vs containerized deployment
For each decision, explain:
Requirement → Options → Trade-offs → Decision → Cost → Performance → Security → Business Impact
How to apply
You will be taken to the employer’s site to complete your application.
Apply for this job (opens in a new tab)