
In Part 1 of this series, we built an agentic AI system that could search notes, create content, and manage topics — all through natural conversation. The agent decided when and how to use its tools, which is what makes it agentic rather than just a chatbot.
But those tools only returned text. In a real application, users will ask the agent to produce files — export data as Excel, generate reports, convert formats — and to understand images and uploaded documents. That’s what we’ll build in this article.
We’ll cover four practical patterns:
- A file generation tool that converts data to Excel, CSV, or JSON files
- A side-channel pattern for delivering generated files alongside streamed text responses
- An approval workflow for safe write operations — the AI proposes, the user approves
- Multimodal input handling — making your agent understand images and spreadsheets
All code examples are simplified versions of production tools. You can copy them, adapt them, improve and ship them.
The File Generation Tool
This is the most impactful tool you can give an agent. Once it can generate files, conversations like this become possible:
User: "List all my notification templates and export them as Excel"
Agent: [calls ListTemplates tool → gets data → calls SaveAsFile → Excel generated]
"Here's your templates exported as templates.xlsx (12.4 KB)"
The agent chains tools autonomously— it reads data with one tool, then exports it with another. You don’t hard-code this workflow; the AI decides to do it based on the user’s request.
The Tool Definition
using System.ComponentModel;
using System.Text;
using System.Text.Json;
using DocumentFormat.OpenXml;
using DocumentFormat.OpenXml.Packaging;
using DocumentFormat.OpenXml.Spreadsheet;
public class FileTools(AgentContext context)
{
[Description("Save content as a downloadable file. Supports .json, .csv, " +
".txt, .xlsx (Excel). For .xlsx, provide content as a JSON array of objects.")]
public async Task<string> SaveAsFileAsync(
[Description("The content to save. For .xlsx, provide a JSON array of objects.")]
string content,
[Description("File name with extension, e.g. 'report.xlsx'")]
string fileName,
CancellationToken ct = default)
{
var extension = Path.GetExtension(fileName).ToLowerInvariant();
var (data, contentType) = extension switch
{
".xlsx" => (ConvertJsonToExcel(content),
"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet"),
".csv" => (Encoding.UTF8.GetBytes(content), "text/csv"),
".json" => (Encoding.UTF8.GetBytes(content), "application/json"),
_ => (Encoding.UTF8.GetBytes(content), "text/plain")
};
// Save the file (replace with your storage: Azure Blob, S3, local disk, etc.)
var fileId = await SaveToStorageAsync(data, fileName, contentType, ct);
// Push metadata to side-channel (more on this in the next section)
context.GeneratedFiles.Add(new GeneratedFile
{
FileId = fileId,
FileName = fileName,
ContentType = contentType,
SizeBytes = data.Length
});
return $"Saved {fileName} ({FormatSize(data.Length)}). The file is attached for download.";
}
}
The “[Description]” attribute is critical — it tells the AI when and how to use the tool. Notice we explicitly mention the “.xlsx” format expects a JSON array. This guides the model to prepare the data correctly before calling the tool.
The switch expression handles format routing cleanly. For CSV, JSON, and text, the content goes straight through. For Excel, we do the heavy lifting.
Converting JSON to Excel
When the AI calls “SaveAsFileAsync” with an “.xlsx” extension, it passes a JSON array that looks something like:
[
{ "name": "Order Confirmation", "locale": "en", "lastModified": "2025-01-15" },
{ "name": "Shipping Update", "locale": "en", "lastModified": "2025-02-01" },
{ "name": "Return Label", "locale": "fi", "lastModified": "2025-01-20" }
]
Our “ConvertJsonToExcel” method turns this into a proper “.xlsx” file with typed cells:
private static byte[] ConvertJsonToExcel(string jsonContent)
{
var rows = JsonSerializer.Deserialize<List<Dictionary<string, JsonElement>>>(jsonContent)
?? throw new InvalidOperationException("Content must be a JSON array of objects.");
// Flatten nested objects: { "order": { "number": 1191 } } → { "order.number": 1191 }
var flatRows = rows.Select(FlattenRow).ToList();
var ms = new MemoryStream();
using (var doc = SpreadsheetDocument.Create(ms, SpreadsheetDocumentType.Workbook))
{
var workbookPart = doc.AddWorkbookPart();
workbookPart.Workbook = new Workbook();
var worksheetPart = workbookPart.AddNewPart<WorksheetPart>();
worksheetPart.Worksheet = new Worksheet(new SheetData());
var sheets = workbookPart.Workbook.AppendChild(new Sheets());
sheets.Append(new Sheet
{
Id = workbookPart.GetIdOfPart(worksheetPart),
SheetId = 1,
Name = "Data"
});
var sheetData = worksheetPart.Worksheet.GetFirstChild<SheetData>()!;
// Collect all unique column headers across all rows
var headers = flatRows.SelectMany(r => r.Keys).Distinct().ToList();
// Write header row
var headerRow = new Row { RowIndex = 1 };
for (int i = 0; i < headers.Count; i++)
{
headerRow.Append(new Cell
{
CellReference = GetColumnName(i) + "1",
DataType = CellValues.String,
CellValue = new CellValue(headers[i])
});
}
sheetData.Append(headerRow);
// Write data rows with type-aware cells
uint rowIndex = 2;
foreach (var row in flatRows)
{
var dataRow = new Row { RowIndex = rowIndex };
for (int i = 0; i < headers.Count; i++)
{
if (row.TryGetValue(headers[i], out var value))
dataRow.Append(CreateCell(GetColumnName(i), rowIndex, value));
}
sheetData.Append(dataRow);
rowIndex++;
}
workbookPart.Workbook.Save();
}
return ms.ToArray();
}
Two details worth highlighting:
Flattening nested JSON. Real-world data is often nested. The flattener converts `{ “order”: { “number”: 1191 } }` into `{ “order.number”: 1191 }` so every value gets its own Excel column:
private static Dictionary<string, JsonElement> FlattenRow(Dictionary<string, JsonElement> row)
{
var flat = new Dictionary<string, JsonElement>();
foreach (var (key, value) in row)
FlattenValue(key, value, flat);
return flat;
}
private static void FlattenValue(string prefix, JsonElement value,
Dictionary<string, JsonElement> flat)
{
if (value.ValueKind == JsonValueKind.Object)
{
foreach (var prop in value.EnumerateObject())
FlattenValue($"{prefix}.{prop.Name}", prop.Value, flat);
}
else
{
flat[prefix] = value;
}
}
Type-aware cell creation. Excel distinguishes between numbers, booleans, and strings. If you write everything as strings, numeric sorting and formulas break. This helper preserves types:
private static Cell CreateCell(string column, uint row, JsonElement value)
{
var cellRef = $"{column}{row}";
return value.ValueKind switch
{
JsonValueKind.Number when value.TryGetInt32(out var intVal) =>
new Cell { CellReference = cellRef, DataType = CellValues.Number,
CellValue = new CellValue(intVal) },
JsonValueKind.Number when value.TryGetDecimal(out var decVal) =>
new Cell { CellReference = cellRef, DataType = CellValues.Number,
CellValue = new CellValue(decVal) },
JsonValueKind.True or JsonValueKind.False =>
new Cell { CellReference = cellRef, DataType = CellValues.String,
CellValue = new CellValue(value.GetBoolean().ToString()) },
JsonValueKind.Null or JsonValueKind.Undefined =>
new Cell { CellReference = cellRef },
_ => new Cell { CellReference = cellRef, DataType = CellValues.String,
CellValue = new CellValue(value.ToString()) }
};
}
And the column name helper that handles A through Z, then AA, AB, etc.:
private static string GetColumnName(int index)
{
var name = "";
index++;
while (index > 0)
{
index--;
name = (char)('A' + index % 26) + name;
index /= 26;
}
return name;
}
NuGet package: “DocumentFormat.OpenXml”
The Side-Channel Pattern
Here’s a problem you’ll hit immediately: the tool returns a string to the AI (“Saved report.xlsx”), which the AI includes in its text response. But the actual file bytes need to reach the UI as a downloadable attachment, not as chat text. The AI never sees the file content — it just confirms the save.
The solution is a side-channel on the agent context. The tool pushes file metadata to a concurrent collection during execution. The streaming loop drains it after each tool call and sends it to the client as a separate stream chunk.
The Context
using System.Collections.Concurrent;
public class AgentContext
{
public string UserId { get; set; } = "";
public string? UserName { get; set; }
/// <summary>
/// Side-channel for files generated by tools (Excel exports, JSON files, etc.)
/// </summary>
public ConcurrentBag<GeneratedFile> GeneratedFiles { get; } = new();
/// <summary>
/// Side-channel for approval requests from write tools.
/// Tools push here instead of executing destructive operations directly.
/// </summary>
public ConcurrentBag<PendingApproval> PendingApprovals { get; } = new();
}
public class GeneratedFile
{
public string FileId { get; set; } = "";
public string FileName { get; set; } = "";
public string ContentType { get; set; } = "";
public long SizeBytes { get; set; }
}
public class PendingApproval
{
public Guid Id { get; set; } = Guid.NewGuid();
public string ToolName { get; set; } = "";
public string ActionType { get; set; } = "";
public string Summary { get; set; } = "";
public string Details { get; set; } = "{}"; // JSON payload for the handler
}
We use ConcurrentBag because tool execution can be concurrent — the framework might invoke multiple tools in parallel. Both bags follow the same pattern: tools push, the streaming loop drains.
Draining in the Streaming Loop
Inside your ChatStreamingAsync method, after processing each FunctionResultContent, drain the side-channel:
await foreach (var update in agent.GetStreamingResponseAsync(messages, chatOptions, ct))
{
// Stream text chunks to the client
if (!string.IsNullOrEmpty(update.Text))
{
yield return new ChatStreamChunk { Text = update.Text };
}
if (update.Contents != null)
{
foreach (var content in update.Contents)
{
// Track tool calls for the UI
if (content is FunctionCallContent functionCall)
{
yield return new ChatStreamChunk
{
ToolCall = new ToolCallInfo
{
Name = functionCall.Name,
Arguments = JsonSerializer.Serialize(functionCall.Arguments)
}
};
}
// After each tool result, drain both side-channels
if (content is FunctionResultContent)
{
while (context.GeneratedFiles.TryTake(out var file))
yield return new ChatStreamChunk { FileAttachment = file };
while (context.PendingApprovals.TryTake(out var approval))
yield return new ChatStreamChunk { ApprovalRequest = approval };
}
}
}
}
The client receives four types of stream chunks interleaved:
- Text — the AI’s response, streamed word by word
- ToolCall — metadata about which tool was invoked (useful for UI indicators)
- FileAttachment— a generated file ready for download
- ApprovalRequest — a write operation waiting for the user to approve or reject
The tool communicates with the AI via its return value, and with the UI via the side-channel. This same pattern works for any tool side effect — files, approvals, notifications.
The Approval Workflow: Safe Write Operations
Read tools are safe — listing templates or fetching settings can’t break anything. But what about tools that modify data? Cancelling a return order, updating a product, or changing a schedule are operations you don’t want the AI to execute without human confirmation.
The approval pattern solves this: write tools don’t execute directly. Instead, they validate the request, push an approval to the side-channel, and tell the AI “approval required.” The user sees an approve/reject prompt in the chat UI.
Write Tools Push Approvals
Here’s a simplified tool that cancels a return order — but only after the user approves:
[Description("Cancel a return order. Requires user approval before execution.")]
public async Task<string> CancelReturnAsync(
[Description("The return order ID")] Guid returnId,
[Description("Reason for cancellation")] string? reason = null,
CancellationToken ct = default)
{
// Validate first — fail fast before asking for approval
var returnOrder = await returnService.GetAsync(returnId, context.ChannelId);
if (returnOrder == null)
return "Return not found.";
if (returnOrder.Status is "Cancelled" or "Completed")
return $"Cannot cancel — return is already '{returnOrder.Status}'.";
// Don't execute — push to approval side-channel instead
context.PendingApprovals.Add(new PendingApproval
{
ToolName = "CancelReturn",
ActionType = "return_cancel",
Summary = $"Cancel return {returnOrder.ReturnId} (status: {returnOrder.Status})",
Details = JsonSerializer.Serialize(new { returnId, reason })
});
return $"[APPROVAL_REQUIRED] Cancel return {returnOrder.ReturnId}. " +
"Please approve or reject this action.";
}The tool returns “[APPROVAL_REQUIRED]” to the AI, which includes this in its response — something like “I’d like to cancel return #1234. Please approve this action.” Meanwhile, the PendingApproval travels through the side-channel to the UI as a structured approval card.
The Details field is a JSON payload containing everything the handler needs to execute later. The tool validates now, but execution happens after approval.
Approval Handlers Execute the Action
When the user clicks “Approve,” a handler executes the actual operation. Each action type has its own handler:
public interface IApprovalActionHandler
{
string ActionType { get; }
Task<(bool Success, string Message)> ExecuteAsync(
AgentContextBase context, JsonElement details, CancellationToken ct);
}
public class ReturnCancelHandler(IReturnService returnService) : IApprovalActionHandler
{
public string ActionType => "return_cancel";
public async Task<(bool Success, string Message)> ExecuteAsync(
AgentContextBase context, JsonElement details, CancellationToken ct)
{
var returnId = details.GetProperty("returnId").GetGuid();
if (details.TryGetProperty("reason", out var r) && r.ValueKind == JsonValueKind.String)
await returnService.UpdateAsync(returnId, new { Result = r.GetString() });
var result = await returnService.CancelReturnOrderAsync(returnId, ct);
return (true, $"Return {result.ReturnId} has been cancelled.");
}
}
A dispatcher routes approvals to the right handler:
public class ApprovalExecutionService(IEnumerable<IApprovalActionHandler> handlers)
{
public async Task<(bool Success, string Message)> ExecuteAsync(
AgentContextBase context, ApprovalRequest approval, CancellationToken ct)
{
var handler = handlers.FirstOrDefault(h => h.ActionType == approval.ActionType);
if (handler == null)
return (false, $"Unknown action type: {approval.ActionType}");
var details = JsonSerializer.Deserialize<JsonElement>(approval.Details);
return await handler.ExecuteAsync(context, details, ct);
}
}
Register handlers per service — each service defines handlers for its own domain:
// In your service registration
services.AddScoped<IApprovalActionHandler, ReturnCancelHandler>();
services.AddScoped<IApprovalActionHandler, ReturnApproveHandler>();
services.AddScoped<IApprovalActionHandler, ProductUpdateHandler>();
The Resolution Endpoint
The frontend calls a simple endpoint to approve or reject:
[HttpPut("conversations/{conversationId}/approvals/{approvalId}")]
public async Task<IActionResult> UpdateApprovalStatus(
Guid conversationId, Guid approvalId,
[FromBody] UpdateApprovalStatusDto dto)
{
// dto.Status = "approved" or "rejected"
// dto.RejectionReason = "..." (optional, for rejections)
if (dto.Status == "approved")
{
var (success, message) = await executionService.ExecuteAsync(context, approval, ct);
dto.ExecutionResult = message;
}
await chatService.UpdateApprovalStatusAsync(
conversationId, approvalId, dto.Status,
dto.ExecutionResult, dto.RejectionReason);
return Ok();
}The full audit trail is preserved: who approved, when, and what the execution result was.
Why This Pattern?
The AI decides what to do. The human decides whether to do it. This gives you the best of both worlds — the agent handles the complex reasoning (“which return should be cancelled, what’s the current status, is it eligible?”), while the user retains control over destructive operations. The approval details are stored as JSON, so you have a complete audit trail of every action the AI proposed and whether it was approved or rejected. There is no way AI could do actions by itself as it does not know how actions are done. There can be an approve all (10) button in the frontend, so approving several actions is not an issue for the user either.
Multimodal Input: Images and File Uploads
So far our tools produce output. But what about input? Users want to send screenshots, upload spreadsheets, and ask the agent to analyse them. This requires multimodal messages — a single user message containing both text and binary content.
Building Multimodal Messages
The Microsoft.Extensions.AI library supports multimodal content natively through AIContent types. When a user sends a message with attachments, we build the ChatMessage like this:
private static ChatMessage BuildUserMessage(
string text, List<FileAttachment>? attachments)
{
if (attachments == null || attachments.Count == 0)
return new ChatMessage(ChatRole.User, text);
var contents = new List<AIContent> { new TextContent(text) };
foreach (var att in attachments)
{
if (att.ContentType.StartsWith("image/"))
{
// Images go as binary data — the LLM "sees" them
contents.Add(new DataContent(att.Content, att.ContentType));
}
else
{
// Text files, CSVs, etc. — extract text and append
var extracted = ExtractText(att);
contents.Add(new TextContent(
$"[Attached file: {att.FileName}]\n{extracted}"));
}
}
return new ChatMessage(ChatRole.User, contents);
}
Images are sent as DataContent with their MIME type — the vision-capable model receives the raw image bytes and can analyse them visually. Text-based files (CSV, JSON, XML) are extracted as text and appended to the message.
Scaling Images for the LLM
Large images waste tokens and slow down responses. Before sending an image to the model, scale it down if it exceeds a reasonable threshold:
using SkiaSharp;
private const int ImageSizeThreshold = 512 * 1024; // 512 KB
private const int MaxImageDimension = 1024;
private const int JpegQuality = 85;
private static byte[] ScaleImage(byte[] imageBytes)
{
using var original = SKBitmap.Decode(imageBytes);
if (original == null)
return imageBytes; // Can't decode — return as-is
var width = original.Width;
var height = original.Height;
if (width <= MaxImageDimension && height <= MaxImageDimension)
{
// Still re-encode as JPEG for consistent size
using var img = SKImage.FromBitmap(original);
using var data = img.Encode(SKEncodedImageFormat.Jpeg, JpegQuality);
return data.ToArray();
}
// Scale down maintaining aspect ratio
float scale = Math.Min(
(float)MaxImageDimension / width,
(float)MaxImageDimension / height);
var newWidth = (int)(width * scale);
var newHeight = (int)(height * scale);
using var resized = original.Resize(
new SKImageInfo(newWidth, newHeight), SKSamplingOptions.Default);
using var image = SKImage.FromBitmap(resized);
using var encoded = image.Encode(SKEncodedImageFormat.Jpeg, JpegQuality);
return encoded.ToArray();
}
This caps images at 1024px on the longest side and re-encodes as JPEG at 85% quality. A 4MB phone photo becomes a ~150KB image that the model processes just as well.
NuGet package: “SkiaSharp”
Reading Excel Files as Input
When a user uploads a spreadsheet, the model can’t read binary Excel formats directly. We convert it to tab-separated text:
using ExcelDataReader;
private static string ReadExcelAsText(byte[] content)
{
Encoding.RegisterProvider(CodePagesEncodingProvider.Instance);
using var stream = new MemoryStream(content);
using var reader = ExcelReaderFactory.CreateReader(stream);
var sb = new StringBuilder();
var rowCount = 0;
do
{
if (reader.Name is not null)
sb.AppendLine($"--- Sheet: {reader.Name} ---");
while (reader.Read())
{
for (var col = 0; col < reader.FieldCount; col++)
{
if (col > 0) sb.Append('\t');
sb.Append(reader.GetValue(col)?.ToString() ?? "");
}
sb.AppendLine();
if (++rowCount >= 2000)
{
sb.AppendLine("[... truncated, sheet has more rows]");
break;
}
}
} while (reader.NextResult());
return sb.ToString();
}
The 2000-row limit prevents token explosion on large spreadsheets. The AI receives clean tabular text it can reason about, summarise, or transform.
NuGet package: “ExcelDataReader”
What This Enables
With these pieces in place, your agent handles flows like:
User: [attaches screenshot] "What does this error mean?"
→ Image sent as DataContent → AI analyzes it visually → explains the error
User: [attaches orders.xlsx] "Summarize these orders by status"
→ Excel converted to text → AI reads, groups, summarizes
User: [attaches orders.xlsx] "Convert this to JSON and save it"
→ Excel read → AI transforms the data → calls SaveAsFile → JSON file generated
The last example is particularly powerful: the agent chains input processing (read Excel) with output generation (write JSON file), all in a single conversational turn.
Wrapping Up
Four patterns, four levels of capability:
- File generation turns your agent from a conversationalist into a productivity tool. Users get downloadable Excel, CSV, and JSON files from natural language requests.
- The side-channel pattern solves the delivery problem — tools communicate file metadata and approval requests to the UI without polluting the AI conversation.
- The approval workflow makes write operations safe. The AI handles the reasoning; the human retains control over destructive actions, with a full audit trail.
- Multimodal input lets users send images and spreadsheets, making the agent useful for analysis and data transformation tasks.
The tools themselves are straightforward — the real insight is how they compose. An agent with input, output, and approval tools can do things like “read this Excel, filter rows where status is pending, export the result, and cancel the overdue returns” — all from a single conversation, with the AI deciding the execution plan and the user approving each write operation.
In the next article, we’ll build a Vue.js frontend with real-time streaming, file download UI, approval cards, and tool execution indicators — connecting everything to a chat interface your users will actually enjoy using.
The code in this article is simplified from production tools being developed for the OGOship logistics platform. Full source uses dependency injection, Azure Blob Storage, and generic type parameters for multi-service reuse.
NuGet packages used:
- DocumentFormat.OpenXml— Excel file generation
- SkiaSharp— Image scaling and re-encoding
- ExcelDataReader— Reading uploaded Excel files
- Microsoft.Extensions.AI— AI abstractions (“IChatClient”, “AIContent”, “DataContent”)
Edit: Part 3 Vue.js implementation is now available here.