Merge pull request #1012 from dotnet/trim-technology-selection-889

Trim technology-selection skill cost via gated progressive disclosure
This commit is contained in:
Abhitej John
2026-08-17 13:40:10 -07:00
committed by GitHub
9 changed files with 359 additions and 346 deletions
@@ -6,350 +6,101 @@ license: MIT
# .NET AI and Machine Learning
## Inputs
| Input | Required | Description |
|-------|----------|-------------|
| Task description | Yes | What the AI/ML feature should accomplish (e.g., "classify support tickets", "summarize documents") |
| Data description | Yes | Type and shape of input data (structured/tabular, unstructured text, images, mixed) |
| Deployment constraints | No | Cloud vs. local, latency SLO, cost budget, offline requirements |
| Existing project context | No | Current .csproj, existing packages, target framework |
## Workflow
### Step 1: Classify the task using the decision tree
Evaluate the developer's task against this decision tree and select the appropriate technology. State which branch applies and why.
| Task type | Technology | Rationale |
|-----------|-----------|-----------|
| Structured/tabular data: classification, regression, clustering, anomaly detection, recommendation | **ML.NET** (`Microsoft.ML`) | Reproducible (given a fixed seed and dataset), no cloud dependency, purpose-built models for these tasks |
| Natural language understanding, generation, summarization, reasoning over unstructured text (single prompt → response, no tool calling) | **LLM via Microsoft.Extensions.AI** (`IChatClient`) | Requires language model capabilities beyond pattern matching; no orchestration needed |
| Agentic workflows: tool/function calling, multi-step reasoning, agent loops, multi-agent collaboration | **Microsoft Agent Framework** (`Microsoft.Agents.AI`) built on top of **Microsoft.Extensions.AI** | Requires orchestration, tool dispatch, iteration control, and guardrails that `IChatClient` alone does not provide |
| Building GitHub Copilot extensions, custom agents, or developer workflow tools | **GitHub Copilot SDK** (`GitHub.Copilot.SDK`) | Integrates with the Copilot agent runtime for IDE and CLI extensibility |
| Running a pre-trained or fine-tuned custom model in production | **ONNX Runtime** (`Microsoft.ML.OnnxRuntime`) | Hardware-accelerated inference, model-format agnostic |
| Local/offline LLM inference with no cloud dependency | **OllamaSharp** with local [AI models supported by Ollama](https://ollama.com/search) | Privacy-sensitive, air-gapped, or cost-constrained scenarios |
| Semantic search, RAG, or embedding storage | **Microsoft.Extensions.VectorData.Abstractions** + a vector database provider (e.g., Azure AI Search, Milvus, MongoDB, pgvector, Pinecone, Qdrant, Redis, SQL) | Provider-agnostic abstractions for vector similarity search; pair with a database-specific connector package (many are moving to community toolkits) |
| Ingesting, chunking, and loading documents into a vector store | **Microsoft.Extensions.AI.DataIngestion** (preview) + **Microsoft.Extensions.VectorData.Abstractions** (MEVD) | Handles document parsing, text chunking, embedding generation, and upserting into a vector database; pairs with Microsoft.Extensions.VectorData.Abstractions |
| Both structured ML predictions AND natural language reasoning | **Hybrid**: ML.NET for predictions + LLM for reasoning layer | Keep loosely coupled; ML.NET handles reproducible scoring, LLM adds explanation |
**Critical rule:** Do NOT use an LLM for tasks that ML.NET handles well (classification on tabular data, regression, clustering). LLMs are slower, more expensive, and non-deterministic for these tasks.
### Step 1b: Select the correct library layer
After identifying the task type, select the right library layer. These libraries form a stack — each builds on the one below it. Using the wrong layer is a major source of non-deterministic agent behavior.
| Layer | Library | NuGet package | Use when |
|-------|---------|---------------|----------|
| **Abstraction** | Microsoft.Extensions.AI (MEAI) | `Microsoft.Extensions.AI` | You need a provider-agnostic interface for chat, embeddings, or tool calling. This is the foundation — always include it. Use `IChatClient` directly **only** for simple prompt-in/response-out scenarios with no tool calling or agentic loops. If the task involves tools, agents, or multi-step reasoning, you must add the Orchestration layer above. |
| **Provider SDK** | OpenAI, Azure.AI.OpenAI, Azure.AI.Inference, OllamaSharp | `OpenAI`, `Azure.AI.OpenAI`, `Azure.AI.Inference`, `OllamaSharp` | You need a concrete LLM provider implementation. These wire into MEAI via `AddChatClient`. Use `OpenAI` for direct OpenAI access, `Azure.AI.OpenAI` for Azure OpenAI, `Azure.AI.Inference` for Azure AI Foundry / GitHub Models, or `OllamaSharp` for local Ollama. Use directly only if you need provider-specific features not exposed through MEAI. |
| **Orchestration** | Microsoft Agent Framework | `Microsoft.Agents.AI` (prerelease) | The task involves tool/function calling, agentic loops, multi-step reasoning, multi-agent coordination, durable context, or graph-based workflows. **This is required whenever the scenario involves agents or tools — do not hand-roll tool dispatch loops with `IChatClient`.** Builds on top of MEAI. **Note:** This package is currently prerelease — use `dotnet add package Microsoft.Agents.AI --prerelease` to install it. |
| **Copilot integration** | GitHub Copilot SDK | `GitHub.Copilot.SDK` | You are building extensions or tools that integrate with the GitHub Copilot runtime — custom agents, IDE extensions, or developer workflow automation that leverages the Copilot agent platform. |
#### Decision rules for library selection
1. **Start with MEAI.** Every AI integration begins with `Microsoft.Extensions.AI` for the `IChatClient` / `IEmbeddingGenerator` abstractions. This ensures provider-swappability and testability.
2. **Add a provider SDK** (`OpenAI`, `Azure.AI.OpenAI`) as the concrete implementation behind MEAI. Do not call the provider SDK directly in business logic — always go through the MEAI abstraction.
3. **Use Agent Framework (`Microsoft.Agents.AI`) for any task that involves tools or agents.** If the task is a single prompt → response with no tool calling, MEAI is sufficient. **You MUST use `Microsoft.Agents.AI`** when any of these apply:
- Tool/function calling (agent decides which tools to invoke)
- Multi-step reasoning with state carried across turns
- Agentic loops that iterate until a goal is met
- Multi-agent collaboration with handoff protocols
- Graph-based or durable workflows
Do **not** implement these patterns by hand with `IChatClient` — the Agent Framework provides iteration limits, observability, and tool dispatch that are error-prone to reimplement.
4. **Add Copilot SDK only when building Copilot extensions.** Use `GitHub.Copilot.SDK` when the goal is to build a custom agent or tool that runs inside the GitHub Copilot platform (CLI, IDE, or Copilot Chat). This is not a general-purpose LLM orchestration library — it is specifically for Copilot extensibility.
5. **Never skip layers.** Do not use Agent Framework without MEAI underneath. Do not call `HttpClient` to OpenAI alongside MEAI in the same workflow. Each layer depends on the one below it.
### Step 2: Select packages and set up the project
Install only the packages needed for the selected technology branch. Do not mix competing abstractions.
#### Classic ML packages
```xml
<PackageReference Include="Microsoft.ML" Version="4.*" />
<PackageReference Include="Microsoft.ML.AutoML" Version="0.*" />
<!-- Only if custom numerical work is needed: -->
PackageReference Include="System.Numerics.Tensors" Version="10.*"
<PackageReference Include="MathNet.Numerics" Version="5.*" />
<!-- Only for data exploration: -->
<PackageReference Include="Microsoft.Data.Analysis" Version="0.*" />
```
> **Do NOT use** Accord.NET — it is archived and unmaintained.
#### Modern AI packages
```xml
<!-- Always start with the abstraction layer -->
<PackageReference Include="Microsoft.Extensions.AI" Version="9.*" />
<!-- Orchestration (agents, workflows, tools, memory) — prerelease; use dotnet add package Microsoft.Agents.AI --prerelease -->
<PackageReference Include="Microsoft.Agents.AI" Version="1.*-*" />
<!-- Cloud LLM provider (pick one) -->
<PackageReference Include="Azure.AI.OpenAI" Version="2.*" />
<!-- OR -->
<PackageReference Include="OpenAI" Version="2.*" />
<!-- Client-side token counting for cost management -->
<PackageReference Include="Microsoft.ML.Tokenizers" Version="2.*"
<!-- Local LLM inference -->
<PackageReference Include="OllamaSharp" Version="5.*" />
<!-- Custom model inference -->
<PackageReference Include="Microsoft.ML.OnnxRuntime" Version="1.*" />
<!-- Vector store abstraction -->
<PackageReference Include="Microsoft.Extensions.VectorData.Abstractions" Version="9.*" />
<!-- Document ingestion, chunking, and vector store loading (preview) -->
<PackageReference Include="Microsoft.Extensions.AI.DataIngestion" Version="9.*-*" />
<!-- Copilot platform extensibility -->
<PackageReference Include="GitHub.Copilot.SDK" Version="1.*" />
```
> **Stack coherence rule:** Never mix raw SDK calls (`HttpClient` to OpenAI) with `Microsoft.Extensions.AI`, Microsoft Agent Framework, or Copilot SDK in the same workflow. Pick one abstraction layer per workflow boundary and commit to it. See Step 1b for the layering rules.
#### Register services with dependency injection
All AI/ML services must be registered via DI. Never instantiate clients directly in business logic.
```csharp
// Configuration via IOptions<T>
services.Configure<AiOptions>(configuration.GetSection("AI"));
// Register the AI client through the abstraction
services.AddChatClient(builder => builder
.UseOpenAIChatClient("gpt-4o-mini-2024-07-18"));
```
### Step 3: Implement with guardrails
Apply the guardrails for the selected technology branch. Every generated implementation must follow these rules.
#### Classic ML guardrails
1. **Reproducibility**: Always set a random seed in the ML context:
```csharp
var mlContext = new MLContext(seed: 42);
```
2. **Data splitting**: Always split into train/test (and optionally validation). Never evaluate on training data:
```csharp
var split = mlContext.Data.TrainTestSplit(data, testFraction: 0.2);
```
3. **Metrics logging**: Always compute and log evaluation metrics appropriate to the task:
```csharp
var metrics = mlContext.BinaryClassification.Evaluate(predictions);
logger.LogInformation("AUC: {Auc:F4}, F1: {F1:F4}", metrics.AreaUnderRocCurve, metrics.F1Score);
```
4. **AutoML first**: Prefer `mlContext.Auto()` for initial model selection, then refine manually.
5. **PredictionEngine pooling**: In ASP.NET Core, always use the pooled prediction engine — never a singleton:
```csharp
services.AddPredictionEnginePool<ModelInput, ModelOutput>()
.FromFile(modelPath);
```
#### LLM integration guardrails
1. **Temperature**: Always set explicitly. Use `0` for factual/deterministic tasks:
```csharp
var options = new ChatOptions
{
Temperature = 0f,
MaxOutputTokens = 1024,
};
```
2. **Structured output**: Always parse LLM output into strongly-typed objects with fallback handling:
```csharp
var result = await chatClient.GetResponseAsync<MySchema>(prompt, options, cancellationToken);
```
3. **Retry logic**: Always implement retry with exponential backoff:
```csharp
services.AddChatClient(builder => builder
.UseOpenAIChatClient(modelId)
.Use(new RetryingChatClient(maxRetries: 3)));
```
4. **Cost control**: Always estimate and log token usage. Use Microsoft.ML.Tokenizers to count tokens client-side before sending requests so you can enforce budgets proactively. Choose the smallest model tier that meets quality requirements (e.g., gpt-4o-mini before gpt-4o).
5. **Secret management**: Never hardcode API keys. Use Azure Key Vault, user-secrets, or environment variables:
```csharp
var apiKey = configuration["AI:ApiKey"]
?? throw new InvalidOperationException("AI:ApiKey not configured");
```
6. **Model version pinning**: Specify exact model versions to reduce behavioral drift:
```csharp
// Pin to a specific dated version, not just "gpt-4o"
var modelId = "gpt-4o-2024-08-06";
```
#### Agentic workflow guardrails
0. **Use `Microsoft.Agents.AI` for all agentic workflows.** Do not implement tool dispatch loops or multi-step agent reasoning by hand with `IChatClient`. The Agent Framework provides `ChatClientAgent` (or `AgentWorker`) which handles the tool call → result → re-prompt cycle with built-in guardrails. All rules below assume you are using `Microsoft.Agents.AI`.
1. **Iteration limits**: Always cap agentic loops to prevent runaway execution:
```csharp
var settings = new AgentInvokeOptions
{
MaximumIterations = 10,
};
```
2. **Cost ceiling**: Implement a token budget per execution and terminate when reached. Use Microsoft.ML.Tokenizers to count prompt and completion tokens locally and compare against the budget before each iteration.
3. **Observability**: Log non-sensitive metadata for every agent step. Never log raw `message.Content` — it may contain user prompts, tool outputs, secrets, or PII that persist in plaintext in central logging systems:
```csharp
await foreach (var message in agent.InvokeStreamingAsync(history, settings))
{
logger.LogDebug("Agent step: Role={Role}, ContentLength={Length}",
message.Role, message.Content?.Length ?? 0);
}
```
4. **Tool schemas**: Define explicit tool/function schemas with descriptions. Never rely on implicit tool discovery.
5. **Simplicity preference**: Prefer single-agent with tools over multi-agent unless the task genuinely requires agent collaboration.
#### RAG guardrails
1. **Embedding caching**: Never re-embed the same content on every query. Cache embeddings in the vector store.
2. **Chunking strategy**: Use semantic chunking (split on paragraph/section boundaries) over fixed-size chunking. Ensure chunks have enough context to be useful on their own.
3. **Relevance thresholds**: Do not inject low-relevance chunks into context. Set a minimum similarity score:
```csharp
var results = await vectorStore.SearchAsync(query, new VectorSearchOptions
{
Top = 5,
MinimumScore = 0.75f,
});
```
4. **Source attribution**: Track which chunks contributed to the final response. Include source references in the output.
5. **Batch embeddings**: Batch embedding API calls where possible to reduce latency and cost.
### Step 4: Handle non-determinism
When the solution involves LLM calls or agentic workflows, explicitly address non-determinism:
1. **Acknowledge it**: Inform the developer that LLM outputs are non-deterministic even at temperature 0 (due to batching, quantization, and model updates).
2. **Validate outputs**: Implement schema validation and content assertion checks on every LLM response.
3. **Graceful degradation**: Design a fallback path for when the LLM returns unexpected, malformed, or empty output:
```csharp
var response = await chatClient.GetResponseAsync<ClassificationResult>(prompt, options);
if (response is null || !response.IsValid())
{
logger.LogWarning("LLM returned invalid response, falling back to rule-based classifier");
return ruleBasedClassifier.Classify(input);
}
```
4. **Evaluation harness**: For any prompt that will be iterated on, recommend creating a golden dataset and evaluation scaffold to measure prompt quality over time.
5. **Model version pinning**: Pin to specific dated model versions (e.g., `gpt-4o-2024-08-06`) to reduce drift between deployments.
### Step 5: Apply performance and cost controls
1. **Connection pooling**: Use `IHttpClientFactory` and DI-managed clients for all external services.
2. **Response caching**: Cache repeated or similar queries. Consider semantic caching for LLM responses where appropriate.
3. **Streaming**: Use `IAsyncEnumerable` for LLM responses in user-facing scenarios to reduce time-to-first-token:
```csharp
await foreach (var update in chatClient.GetStreamingResponseAsync(prompt, options))
{
yield return update.Text;
}
```
4. **Health checks**: Implement health checks for external AI service dependencies:
```csharp
services.AddHealthChecks()
.AddCheck<OpenAIHealthCheck>("openai");
```
5. **ML.NET prediction pooling**: In web applications, always use `PredictionEnginePool<TIn, TOut>`, never a single `PredictionEngine` instance (it is not thread-safe).
### Step 6: Validate the implementation
1. Build the project and verify no warnings:
```bash
dotnet build -c Release -warnaserror
```
2. Run tests, including integration tests that validate AI/ML behavior:
```bash
dotnet test -c Release
```
3. For ML.NET pipelines, verify that evaluation metrics meet the project's quality bar and that the model can be serialized and loaded correctly.
4. For LLM integrations, verify that structured output parsing handles both valid and malformed responses.
5. For RAG pipelines, verify that retrieval returns relevant results and that irrelevant chunks are filtered out.
Pick the right technology first, then deliver **only what the task asks for**. If the task asks for
a plan, comparison, or architecture (or says "do not write code"), produce that — do not scaffold,
build, or run code unprompted.
## Step 1: Classify the task (decision tree)
State which branch applies and why, then choose that technology.
| Task type | Technology | Why |
|-----------|-----------|-----|
| Structured/tabular: classification, regression, clustering, anomaly detection, recommendation | **ML.NET** (`Microsoft.ML`) | Deterministic (fixed seed), no cloud dependency, purpose-built |
| NL understanding, generation, summarization, reasoning (single prompt → response, no tools) | **LLM via Microsoft.Extensions.AI** (`IChatClient`) | Language capability, no orchestration needed |
| Agentic: multi-step tool/function calling, agent loops, multi-agent | **Microsoft Agent Framework** (`Microsoft.Agents.AI`) on **Microsoft.Extensions.AI** | Needs orchestration, tool dispatch, iteration control `IChatClient` lacks |
| GitHub Copilot extensions / custom dev-workflow agents | **GitHub Copilot SDK** (`GitHub.Copilot.SDK`) | Integrates with the Copilot agent runtime |
| Run a pre-trained/custom model in production | **ONNX Runtime** (`Microsoft.ML.OnnxRuntime`) | Hardware-accelerated, format-agnostic inference |
| Local/offline LLM inference | **OllamaSharp** ([Ollama models](https://ollama.com/search)) | Privacy-sensitive, air-gapped, cost-constrained |
| Semantic search, RAG, embedding storage | **Microsoft.Extensions.VectorData.Abstractions** (MEVD) + a provider (Azure AI Search, Milvus, MongoDB, pgvector, Pinecone, Qdrant, Redis, SQL) | Provider-agnostic vector search |
| Ingest, chunk, load documents into a vector store | **Microsoft.Extensions.AI.DataIngestion** (preview) + MEVD | Parses, chunks, embeds, upserts |
| Both structured predictions AND NL reasoning | **Hybrid**: ML.NET scoring + LLM reasoning layer | ML.NET is reproducible; LLM adds explanation |
**Critical rule:** Do NOT use an LLM for tasks ML.NET handles well (tabular classification,
regression, clustering) — LLMs are slower, costlier, and non-deterministic for these.
## Step 1b: Pick the library layer
| Layer | Library | Use when |
|-------|---------|----------|
| **Abstraction** | `Microsoft.Extensions.AI` (MEAI) | Always the foundation. Use `IChatClient` directly for prompt-response and simple, bounded function invocation. |
| **Provider SDK** | `Azure.AI.OpenAI` / `OpenAI` / `Azure.AI.Inference` / `OllamaSharp` | Concrete provider behind MEAI via `AddChatClient`. |
| **Orchestration** | `Microsoft.Agents.AI` (prerelease) | Multi-step tool use, durable agent loops, and multi-agent workflows. |
| **Copilot** | `GitHub.Copilot.SDK` | Building Copilot-platform extensions only. |
Rules: start with MEAI; put the provider behind it via `AddChatClient` (don't call the provider in
business logic); use `Microsoft.Agents.AI` for multi-step or durable agent workflows rather than
hand-rolling an agent loop; never mix a raw `HttpClient`-to-OpenAI call with MEAI in the same
workflow. Do **not** use Accord.NET (archived). For new projects, prefer MEAI and Agent Framework
unless existing Semantic Kernel features or investments are a requirement. Register AI/ML services
via DI; load secrets from user-secrets / env / Key Vault — never hardcode keys.
## Step 2: Cover the branch essentials, then decide depth
Every answer — plan or implementation — must address the guardrails for the selected branch:
- **ML.NET** — `new MLContext(seed: …)` (reproducible); `TrainTestSplit` + evaluate on the held-out
set; report real metrics (MicroAccuracy/MacroAccuracy/LogLoss, AUC/F1, or RMSE/R²); serve with
`PredictionEnginePool<TIn,TOut>` (never a singleton `PredictionEngine`).
- **LLM (MEAI)** — depend on `IChatClient` registered via `AddChatClient` (provider behind it);
set `Temperature` and `MaxOutputTokens` in `ChatOptions`; add retry/timeout
(`RetryingChatClient`/Polly); pin a dated model; load keys from user-secrets / env / Key Vault —
**never hardcode an `sk-…` key**; validate non-deterministic output against a schema with a
fallback.
- **Agentic (Agent Framework)** — orchestrate with `Microsoft.Agents.AI` on `IChatClient` (never a
hand-rolled loop); set `MaximumIterations` and a token/cost ceiling; define each tool with a clear
schema (`AIFunctionFactory.Create`); log each step (never raw sensitive content).
- **RAG / embeddings** — semantic **chunking** (not fixed-size); `IEmbeddingGenerator` and **cache
the embeddings** (don't re-embed per query); store/query with
`Microsoft.Extensions.VectorData.Abstractions` (MEVD) + the provider the user asked for (e.g.
pgvector); filter by a **minimum similarity score**; keep **source attribution** for each answer.
Honor the UI/storage the user specified; use only real, existing NuGet packages.
**Then choose depth:**
- **Plan / comparison / architecture only** (or "do not write code"): answer from this file alone
using the essentials above. **Do NOT open a reference** — the branch essentials here are
sufficient for a selection or plan. For RAG plans, cover chat, ingestion/chunking, embeddings,
vector storage, source attribution, and the requested UI/storage.
- **Writing implementation code**: read the matching reference(s) for packages and implementation
guidance (read only the selected branch; for Hybrid, read both Classic ML.NET and LLM):
- Classic ML.NET → [`references/classic-ml.md`](references/classic-ml.md)
- LLM integration (MEAI) → [`references/llm.md`](references/llm.md)
- Agentic (Agent Framework) → [`references/agentic.md`](references/agentic.md)
- RAG / embeddings / ingestion → [`references/rag.md`](references/rag.md)
- GitHub Copilot extensions → [`references/copilot.md`](references/copilot.md)
- ONNX Runtime inference → [`references/onnx.md`](references/onnx.md)
- Local/offline LLM with Ollama → [`references/ollama.md`](references/ollama.md)
## Validation
- [ ] Technology selection follows the decision tree — LLMs are not used for tasks ML.NET handles
- [ ] All AI/ML services are registered via dependency injection
- [ ] Configuration uses `IOptions<T>` pattern — no hardcoded values
- [ ] API keys are loaded from secure sources — not in source code or committed config files
- [ ] ML.NET pipelines set a random seed and split data for evaluation
- [ ] LLM calls set temperature, max tokens, and retry logic explicitly
- [ ] Agentic workflows have iteration limits and cost ceilings
- [ ] RAG pipelines implement chunking, relevance thresholds, and source attribution
- [ ] Non-deterministic outputs have validation and fallback paths
- [ ] `dotnet build -c Release -warnaserror` completes cleanly
- [ ] Selection follows the decision tree — no LLM for tasks ML.NET handles
- [ ] Only what was asked is produced (plan-only requests get a plan, not code)
- [ ] AI/ML services registered via DI; config via `IOptions<T>`; keys from secure sources
- [ ] Branch guardrails (Step 2 essentials, plus the reference when implementing) are satisfied
- [ ] After implementing, build and run existing tests
## Anti-Patterns to Reject
When reviewing or generating code, flag and redirect the developer if any of these patterns are detected:
| Anti-pattern | Redirect |
|-------------|----------|
| Using an LLM for classification on structured/tabular data | Use ML.NET instead — it is faster, cheaper, and deterministic |
| Calling LLM APIs without retry or timeout logic | Add `RetryingChatClient` or Polly-based retry with exponential backoff |
| Storing API keys in `appsettings.json` committed to source control | Use user-secrets (dev), environment variables, or Azure Key Vault (prod) |
| Using Accord.NET for new projects | Migrate to ML.NET — Accord.NET is archived and unmaintained |
| Building custom neural networks in .NET from scratch | Use a pre-trained model via ONNX Runtime or call an LLM API |
| RAG without chunking strategy or relevance filtering | Implement semantic chunking and set a minimum similarity score threshold |
| Agentic loops without iteration limits or cost ceilings | Add `MaximumIterations` and a token budget ceiling |
| Using MEAI `IChatClient` with raw `HttpClient` calls to the same provider | Pick one abstraction layer and commit to it |
| Implementing tool calling or agentic loops manually with `IChatClient` instead of using `Microsoft.Agents.AI` | Use `Microsoft.Agents.AI` — it provides iteration limits (`MaximumIterations`), built-in tool dispatch, observability hooks, and cost controls. Hand-rolled loops lack these guardrails. |
| Using Agent Framework for a single prompt→response call | Use MEAI `IChatClient` directly — Agent Framework is for multi-step orchestration |
| Using Copilot SDK for general-purpose LLM apps | Copilot SDK is for Copilot platform extensions only — use MEAI + Agent Framework for standalone apps |
| Calling OpenAI SDK directly in business logic instead of through MEAI | Register the provider via `AddChatClient` and depend on `IChatClient` in business code |
| Using `PredictionEngine` as a singleton in ASP.NET Core | Use `PredictionEnginePool<TIn, TOut>` — `PredictionEngine` is not thread-safe |
| Using `Func<ReadOnlySpan<T>>` for delegates with ref struct parameters | Define a custom delegate type — ref structs cannot be generic type arguments |
| Using `Microsoft.SemanticKernel` for new projects | Use `Microsoft.Extensions.AI` + `Microsoft.Agents.AI` — Semantic Kernel is superseded by these newer abstractions for LLM orchestration and tool calling |
## Common Pitfalls
| Pitfall | Solution |
|---------|----------|
| Over-engineering with LLMs | Start with the simplest approach (rules, ML.NET) and add LLM capability only when simpler methods fall short |
| Evaluating ML models on training data | Always use `TrainTestSplit` and report metrics on the held-out test set |
| LLM output drift between deployments | Pin to specific dated model versions (e.g., `gpt-4o-2024-08-06`) |
| Token cost surprises | Set `MaxOutputTokens`, use Microsoft.ML.Tokenizers for accurate client-side token counting, log token counts per request, and alert on budget thresholds |
| Non-reproducible ML training | Set `MLContext(seed: N)` and version your training data alongside the code |
| RAG returning irrelevant context | Set a minimum similarity score and limit the number of injected chunks |
| Cold start latency on ML.NET models | Pre-warm the `PredictionEnginePool` during application startup |
| Microsoft Agent Framework + raw OpenAI SDK in same class | Choose one orchestration layer per workflow boundary |
| LLM for tabular classification | Use **ML.NET** faster, cheaper, deterministic |
| LLM calls without retry/timeout | Add `RetryingChatClient` or Polly retry |
| API keys in committed `appsettings.json` | user-secrets / env / Key Vault |
| Accord.NET, or defaulting to Semantic Kernel without a requirement | ML.NET; prefer MEAI + `Microsoft.Agents.AI` for new work |
| Hand-rolled multi-step tool loops with `IChatClient` | `Microsoft.Agents.AI` (`MaximumIterations`, tool dispatch) |
| Agent Framework for a single prompt→response | `IChatClient` directly |
| Raw `HttpClient`/OpenAI SDK in business logic alongside MEAI | one abstraction layer; depend on `IChatClient` |
| `PredictionEngine` singleton in ASP.NET Core | `PredictionEnginePool<TIn,TOut>` (not thread-safe) |
| RAG without chunking or relevance filtering | semantic chunking + minimum similarity score |
| Building custom neural nets in .NET from scratch | pre-trained via ONNX Runtime or an LLM API |
@@ -0,0 +1,46 @@
# Agentic workflows with Microsoft Agent Framework
Use when the task needs tools/function calling, multi-step reasoning, agent loops, or multiple
agents. Built **on top of** Microsoft.Extensions.AI — never hand-roll a tool loop on `IChatClient`.
## Packages
```xml
<PackageReference Include="Microsoft.Extensions.AI" Version="9.*" />
<PackageReference Include="Microsoft.Agents.AI" Version="1.*-*" /> <!-- prerelease: dotnet add --prerelease -->
<PackageReference Include="Azure.AI.OpenAI" Version="2.*" /> <!-- or another MEAI provider -->
<PackageReference Include="Azure.Identity" Version="1.*" />
```
## Guardrails
1. **Framework, not raw loops** — orchestrate with `Microsoft.Agents.AI`; do not loop raw LLM calls
by hand.
2. **Foundation layer** — build on `Microsoft.Extensions.AI` (`IChatClient`).
3. **Bounded iteration** — set `MaximumIterations` to cap the agent loop and prevent runaway
execution.
4. **Explicit tools** — define each tool/function with a clear schema and description
(`AIFunctionFactory.Create`).
5. **Cost ceiling** — enforce a token budget; stop when exceeded.
6. **Observability** — log each step (tool selected, input, output metadata) — never raw sensitive
content.
7. Prefer a **single agent with tools** over multi-agent unless the task truly needs specialization.
## Minimal shape
```csharp
IChatClient chatClient = new AzureOpenAIClient(new Uri(endpoint), new DefaultAzureCredential())
.GetChatClient("gpt-4o-2024-08-06").AsIChatClient();
AIAgent agent = new ChatClientAgent(chatClient, new ChatClientAgentOptions
{
Instructions = "Research the topic, then summarize findings.",
ChatOptions = new ChatOptions
{
Tools = [AIFunctionFactory.Create(WebSearch), AIFunctionFactory.Create(TakeNote)],
},
});
var runOptions = new ChatClientAgentRunOptions { MaximumIterations = 10 };
var result = await agent.RunAsync("Research the .NET 10 release highlights.", options: runOptions);
```
@@ -0,0 +1,49 @@
# Classic ML with ML.NET
Use for structured/tabular tasks: classification, regression, clustering, anomaly detection,
recommendation. Deterministic, local, no cloud dependency.
## Packages
```xml
<PackageReference Include="Microsoft.ML" Version="4.*" />
<PackageReference Include="Microsoft.ML.AutoML" Version="0.*" /> <!-- optional: model search -->
<PackageReference Include="Microsoft.Extensions.ML" Version="4.*" /> <!-- PredictionEnginePool for ASP.NET Core -->
```
## Guardrails
1. **Reproducible seed** — always construct `MLContext` with a fixed seed.
2. **Held-out evaluation** — split with `TrainTestSplit`, evaluate on the test set, never on training data.
3. **Report real metrics** — multiclass: MicroAccuracy, MacroAccuracy, LogLoss; binary: AUC, F1;
regression: RMSE, R².
4. **Thread-safe serving** — in ASP.NET Core use `PredictionEnginePool<TIn,TOut>`, never a singleton
`PredictionEngine` (it is not thread-safe).
5. Prefer `mlContext.Auto()` (AutoML) for initial trainer/hyperparameter selection.
## Minimal shape
```csharp
var mlContext = new MLContext(seed: 42);
var data = mlContext.Data.LoadFromTextFile<TicketRow>("tickets.csv", hasHeader: true, separatorChar: ',');
var split = mlContext.Data.TrainTestSplit(data, testFraction: 0.2);
var pipeline = mlContext.Transforms.Conversion.MapValueToKey("Label", nameof(TicketRow.Category))
.Append(mlContext.Transforms.Text.FeaturizeText("SubjectF", nameof(TicketRow.Subject)))
.Append(mlContext.Transforms.Text.FeaturizeText("DescriptionF", nameof(TicketRow.Description)))
.Append(mlContext.Transforms.Concatenate("Features", "SubjectF", "DescriptionF", nameof(TicketRow.Priority)))
.Append(mlContext.MulticlassClassification.Trainers.SdcaMaximumEntropy())
.Append(mlContext.Transforms.Conversion.MapKeyToValue("PredictedLabel"));
var model = pipeline.Fit(split.TrainSet);
var metrics = mlContext.MulticlassClassification.Evaluate(model.Transform(split.TestSet));
// log metrics.MicroAccuracy, metrics.MacroAccuracy, metrics.LogLoss
// ASP.NET Core endpoint:
builder.Services.AddPredictionEnginePool<TicketRow, TicketPrediction>().FromFile(modelPath);
// inject PredictionEnginePool<TicketRow, TicketPrediction> and call .Predict(input)
```
**Reject LLMs for these tasks.** If asked to use GPT/an LLM for tabular prediction, redirect to
ML.NET with rationale: faster, cheaper, deterministic, no per-call cost.
@@ -0,0 +1,32 @@
# GitHub Copilot SDK extensions
Use only for custom developer workflows that must run through the GitHub Copilot agent runtime.
Do not use it as a general LLM client.
## Package
```xml
<PackageReference Include="GitHub.Copilot.SDK" Version="0.3.0" />
```
The SDK is pre-1.0. Pin an exact version and review release notes before each upgrade.
## Guardrails
1. Start and reuse one `CopilotClient`; stop it during application shutdown.
2. Create a bounded session for each workflow and dispose the session after use.
3. Set the working directory, model, system message, and permission handler explicitly.
4. Default permission requests to deny when no user is available.
5. Subscribe to session error, usage, and completion events before sending a prompt.
6. Enforce a timeout and cancellation token, and record token usage without sensitive content.
## Minimal shape
```csharp
var client = new CopilotClient(new CopilotClientOptions());
await client.StartAsync();
await using var session = await client.CreateSessionAsync(sessionConfig);
await session.SendAsync(new MessageOptions { Prompt = prompt });
await client.StopAsync();
```
@@ -0,0 +1,41 @@
# LLM integration with Microsoft.Extensions.AI
Use for text generation, summarization, reasoning — single prompt → response, **no tools**. If the
task needs tools/agent loops, use `references/agentic.md` instead.
## Packages
```xml
<PackageReference Include="Microsoft.Extensions.AI" Version="9.*" />
<PackageReference Include="Azure.AI.OpenAI" Version="2.*" /> <!-- or OpenAI / Azure.AI.Inference / OllamaSharp -->
<PackageReference Include="Azure.Identity" Version="1.*" />
<PackageReference Include="Microsoft.ML.Tokenizers" Version="2.*" /> <!-- client-side token budgeting -->
```
## Guardrails
1. **Abstraction, not provider** — depend on `IChatClient`; do not call `Azure.AI.OpenAI` / `OpenAI`
directly in business logic.
2. **DI registration** — register via `AddChatClient`; never `new` a client in business logic.
3. **Explicit options** — set `Temperature` (0 for factual/deterministic tasks) and
`MaxOutputTokens` in `ChatOptions`.
4. **Resilience** — wrap with `RetryingChatClient` (or a Polly pipeline) for retry/timeout.
5. **Pinned model** — use a dated version (e.g. `gpt-4o-2024-08-06`), not an unversioned alias.
6. **Safe secrets** — load keys from user-secrets / env / Key Vault. Never hardcode (`sk-...`) keys.
7. **Non-determinism** — output varies even at temperature 0; validate against a schema with a
graceful fallback (`GetResponseAsync<T>`), and count tokens with `Microsoft.ML.Tokenizers`.
## Minimal shape
```csharp
builder.Services.AddChatClient(sp =>
new AzureOpenAIClient(new Uri(cfg["Ai:Endpoint"]!), new DefaultAzureCredential())
.GetChatClient("gpt-4o-2024-08-06").AsIChatClient()
.AsBuilder()
.Use(inner => new RetryingChatClient(inner, maxRetries: 3))
.Build());
var options = new ChatOptions { Temperature = 0f, MaxOutputTokens = 1024 };
var summary = await chatClient.GetResponseAsync(
[new(ChatRole.System, "Summarize concisely."), new(ChatRole.User, document)], options, ct);
```
@@ -0,0 +1,27 @@
# Local LLM inference with OllamaSharp
Use for local or offline prompt-response work when privacy, air-gapped operation, or cloud cost is
the main constraint. Use the same MEAI abstractions as a hosted provider.
## Packages
```xml
<PackageReference Include="Microsoft.Extensions.AI" Version="9.*" />
<PackageReference Include="OllamaSharp" Version="5.*" />
```
## Guardrails
1. Depend on `IChatClient`; keep OllamaSharp behind the MEAI abstraction.
2. Configure the Ollama endpoint and model name; do not hardcode deployment-specific values.
3. Set `Temperature` and `MaxOutputTokens`, and bound prompt size for local memory limits.
4. Add timeout and cancellation handling because model startup and inference can be slow.
5. Verify that the selected model is present before serving traffic.
6. Measure latency and memory on the target hardware; local does not mean free or fast.
## Minimal shape
```csharp
builder.Services.AddChatClient(
new OllamaApiClient(new Uri(options.Endpoint), options.Model));
```
@@ -0,0 +1,30 @@
# ONNX Runtime inference
Use when a trained model is already available in ONNX format and the .NET application only needs
production inference. Train or convert the model outside the application.
## Package
```xml
<PackageReference Include="Microsoft.ML.OnnxRuntime" Version="1.*" />
```
Use the GPU-specific package only when the deployment target and execution provider require it.
## Guardrails
1. Validate model input names, element types, dimensions, and output names at startup.
2. Create and warm one `InferenceSession` through DI; do not reload the model per request.
3. Normalize and tokenize input exactly as the model expects.
4. Dispose inference results and other native-memory-backed values promptly.
5. Bound input sizes and batch sizes, and measure latency and memory on the deployment hardware.
6. Pin and record the model artifact version with its preprocessing contract.
## Minimal shape
```csharp
builder.Services.AddSingleton(_ => new InferenceSession(modelPath));
var input = NamedOnnxValue.CreateFromTensor("input", tensor);
using var results = session.Run([input]);
```
@@ -0,0 +1,34 @@
# RAG, embeddings, and document ingestion
Use for semantic search and retrieval-augmented Q&A over documents. Two concerns: **ingestion**
(parse → chunk → embed → store) and **query** (embed question → vector search → ground the answer).
## Packages
```xml
<PackageReference Include="Microsoft.Extensions.AI" Version="9.*" /> <!-- IEmbeddingGenerator, IChatClient -->
<PackageReference Include="Microsoft.Extensions.VectorData.Abstractions" Version="9.*" /> <!-- MEVD -->
<PackageReference Include="Microsoft.Extensions.AI.DataIngestion" Version="9.*-*" /> <!-- preview: parse/chunk/embed/upsert -->
<!-- + a vector provider, e.g. pgvector for PostgreSQL, Azure AI Search, Qdrant, Redis -->
```
## Guardrails
1. **Abstractions**`IEmbeddingGenerator` for embeddings,
`Microsoft.Extensions.VectorData.Abstractions` (MEVD) for the store, `IChatClient` for generation.
2. **Semantic chunking** — chunk on paragraph/semantic boundaries, not naive fixed-size cuts.
Use `Microsoft.Extensions.AI.DataIngestion` (or equivalent parse/chunk) for PDFs/markdown.
3. **Relevance threshold** — filter retrieved chunks by a **minimum similarity score**; don't feed
low-scoring noise to the model.
4. **Source attribution** — track which document chunks contributed to each answer.
5. **Cache embeddings** — persist embeddings; never re-embed the corpus on every query. Batch
embedding calls during ingestion.
## Minimal query shape (provider-specific pseudocode)
```csharp
var queryEmbedding = await embeddingGenerator.GenerateAsync(question, ct);
var hits = await SearchProviderAsync(queryEmbedding, top: 5, cancellationToken: ct);
var grounded = hits.Where(h => h.Score >= 0.75); // minimum similarity threshold
// build prompt with the grounded chunks + their source ids for attribution, then IChatClient.GetResponseAsync
```
+10 -7
View File
@@ -70,9 +70,10 @@ stimuli:
- Implements retry logic (e.g., RetryingChatClient or Polly-based)
- Loads API keys from configuration or environment — not hardcoded in source
- Pins to a specific dated model version (e.g., gpt-4o-2024-08-06) rather than an unversioned alias
- name: Reject LLM for tabular classification
prompt: I have a .NET 10 project with a CSV of customer data (Age, Income, Region, PurchaseHistory columns). I want to
predict whether a customer will churn (yes/no).
- name: Select ML.NET for tabular churn prediction
prompt: I have a .NET 10 project and want to add an AI-powered feature that flags customers who are likely to
churn. The provided CSV has Age, Income, Region, PurchaseHistory, and Churned columns. Implement it, and tell me
which technology you chose and why.
environment:
files:
- src: ./fixtures/reject-llm-for-tabular-classification/ChurnPredictor/ChurnPredictor.csproj
@@ -88,10 +89,12 @@ stimuli:
- type: exit-success
- type: prompt
rubric:
- Redirects the user away from using GPT-4/LLM for this tabular classification task
- Recommends ML.NET as the appropriate technology with a clear rationale (faster, cheaper, deterministic)
- Provides an ML.NET-based binary classification implementation instead
- "Follows the decision tree from the skill: structured tabular data → ML.NET"
- Selects ML.NET (Microsoft.ML) for churn prediction — does not use an LLM, GPT, or a generative AI service for
this tabular prediction task
- Explains why that technology fits, with at least one substantive reason grounded in the task (e.g.,
determinism/reproducibility, cost, latency, or suitability for structured tabular data)
- Implements binary classification over the provided CSV — trains on the Churned label using the Age, Income,
Region, and PurchaseHistory columns and produces a churn prediction
- name: Agentic workflow with guardrails
prompt: Build a .NET 10 console app that uses an AI agent to research a topic by searching the web, then summarize
findings. The agent should have access to a web search tool and a note-taking tool.