Skill

Develop Production-Grade LLM Applications

An AI engineer skill for building production LLM applications, advanced RAG systems, and intelligent agents.

Works with openaianthropicollamavllmtgi

78
Spark score
out of 100
Updated 21 days ago
Version 14.1.0
Models
gpt 4oclaude 3 5 sonnetclaude 3 opusllama 3

Add to Favorites

Why it matters

Build and optimize advanced LLM applications, RAG systems, and intelligent agents. Ensure production-readiness with robust architecture, safety, and cost controls.

Outcomes

What it gets done

01

Design and implement LLM-powered features and RAG pipelines.

02

Integrate and manage various LLM models, including open-source options.

03

Develop agent architectures with frameworks like LangChain and CrewAI.

04

Implement monitoring, safety guardrails, and cost optimization strategies.

Install

Add it to your toolbox

Run in your project directory:

curl -fsSL https://spark.entire.vc/get/ag-ai-engineer | bash

Overview

Ai Engineer

This skill is an expert AI engineer persona for production LLM applications: advanced RAG, agent orchestration, vector search, prompt engineering, multimodal AI, and safety/governance controls. Use it when building or improving LLM features, RAG systems, or AI agents. Not for pure data science/ML tasks or work with no AI deployment target.

What it does

An expert AI engineer persona skill for building production-ready LLM applications, advanced RAG systems, and intelligent agents. LLM integration covers OpenAI GPT-4o/o1, Anthropic Claude 4.5/4.1, open-source models (Llama 3.1/3.2, Mixtral, Qwen 2.5, DeepSeek-V2), local deployment (Ollama, vLLM, TGI), production model serving (TorchServe, MLflow, BentoML), multi-model routing, and cost optimization via caching. Advanced RAG covers multi-stage retrieval pipelines, vector databases (Pinecone, Qdrant, Weaviate, Chroma, Milvus, pgvector), embedding models, semantic/recursive/structure-aware chunking, hybrid vector-plus-BM25 search, reranking (Cohere rerank-3, BGE, cross-encoders), query expansion/decomposition/routing, context compression, and advanced patterns like GraphRAG, HyDE, RAG-Fusion, and self-RAG. Agent orchestration covers LangChain/LangGraph, LlamaIndex, CrewAI multi-agent collaboration, AutoGen, the OpenAI Assistants API, short/long-term/episodic agent memory, tool integration (web search, code execution, API/database calls), and agent evaluation. Vector search covers embedding fine-tuning, indexing strategies (HNSW, IVF, LSH), similarity metrics, multi-vector representations, and embedding-drift detection. Prompt engineering covers chain-of-thought/tree-of-thoughts/self-consistency, few-shot optimization, dynamic templates, constitutional AI, prompt versioning/A-B testing, and safety prompting. Production systems cover FastAPI serving, streaming, semantic caching, rate limiting/cost controls, circuit breakers, A/B testing, and observability (LangSmith, Phoenix, Weights & Biases). Multimodal AI covers vision (GPT-4V, Claude Vision, LLaVA, CLIP), audio (Whisper, ElevenLabs), document AI (OCR, LayoutLM), and cross-modal embeddings. AI safety covers content moderation, prompt-injection detection, PII redaction, and bias mitigation.

When to use - and when NOT to

Use it when building or improving LLM features, RAG systems, or AI agents, designing production AI architectures and model integration, optimizing vector search or retrieval pipelines, or implementing AI safety, monitoring, or cost controls. Do not use it for pure data science or traditional ML without LLMs, a quick UI change unrelated to AI features, or tasks with no access to data sources or deployment targets. Always avoid sending sensitive data to external models without approval, and add guardrails for prompt injection, PII, and policy compliance.

Inputs and outputs

Given use cases, constraints, and success metrics, it works through a four-step flow: clarify requirements, design the AI architecture/data flow/model selection, implement with monitoring/safety/cost controls, and validate with tests and a staged rollout plan.

Integrations

Spans the modern AI stack: OpenAI, Anthropic, and open-source model providers; vector databases (Pinecone, Qdrant, Weaviate, Chroma, Milvus, pgvector); agent frameworks (LangChain/LangGraph, LlamaIndex, CrewAI, AutoGen); and observability tooling (LangSmith, Phoenix, Weights & Biases).

Who it's for

AI/ML engineers and platform teams building production LLM applications who need architecture decisions across model selection, RAG design, agent orchestration, multimodal integration, and safety/cost controls, with a staged rollout discipline rather than shipping directly to production.

FAQ

Common questions

Discussion

Questions & comments ยท 0

Sign In Sign in to leave a comment.