← Pekka Ylenius

Building Intelligent Search with Azure Cosmos DB: Full-Text, Vector, and Hybrid Search

31 Jan 2026 · 8 min read · Originally published on Medium

Modern applications demand smarter search capabilities. Users expect to find what they’re looking for even when they don’t use exact keywords. In this article, I’ll show you how to implement full-text search, vector (semantic) search, and hybrid search using Azure Cosmos DB with .NET.

The Use Case: A Notes Application

We’re building a collaborative notes system where users can:

  • Search notes by keywords (traditional full-text search)
  • Find semantically similar notes (vector search)
  • Combine both approaches for best results (hybrid search)

Let’s start with a simplified Note entity:

public class Note
{
[JsonPropertyName("id")]
public string Id { get; set; } = string.Empty;

[JsonPropertyName("pk")]
public string PartitionKey { get; set; } = string.Empty;

[JsonPropertyName("content")]
public string Content { get; set; } = string.Empty;

[JsonPropertyName("createdAt")]
public DateTime CreatedAt { get; set; }

[JsonPropertyName("contentVector")]
public float[]? ContentVector { get; set; }
}

The key field here is “ContentVector” — a float array that stores the semantic embedding of the note’s content.

Part 1: Setting Up Cosmos DB with Vector Search

Step 1: Enable Vector Search in Your Cosmos DB Account

Vector search in Cosmos DB requires the NoSQL API with vector indexing capabilities. Here’s how to enable it:

Via Azure Portal:

  1. Navigate to your Cosmos DB account
  2. Go to Settings → Features
  3. Enable Vector Search for NoSQL API (if not already enabled)

Via Azure CLI:

az cosmosdb update \
--name your-cosmos-account \
--resource-group your-resource-group \
--capabilities EnableNoSQLVectorSearch

Step 2: Create a Container with Vector Index Policy

The container needs a vector indexing policy to enable efficient vector searches. Create the container with the following configuration:

Via Azure Portal:

  1. Create a new container or modify an existing one
  2. In the Indexing Policy section, add the vector index configuration

Via SDK (Recommended):

public async Task CreateContainerWithVectorIndex(CosmosClient client)
{
var database = client.GetDatabase("NotesDb");

var containerProperties = new ContainerProperties
{
Id = "Notes",
PartitionKeyPath = "/pk",

// Vector embedding policy - defines the vector field
VectorEmbeddingPolicy = new VectorEmbeddingPolicy(
new Collection<Embedding>
{
new Embedding
{
Path = "/contentVector",
DataType = VectorDataType.Float32,
Dimensions = 1536, // OpenAI text-embedding-3-small
DistanceFunction = DistanceFunction.Cosine
}
}),

// Indexing policy with vector index
IndexingPolicy = new IndexingPolicy
{
VectorIndexes = new Collection<VectorIndexPath>
{
new VectorIndexPath
{
Path = "/contentVector",
Type = VectorIndexType.QuantizedFlat // or DiskANN for large datasets
}
},
// Standard indexes for text search
IncludedPaths = new Collection<IncludedPath>
{
new IncludedPath { Path = "/*" }
},
ExcludedPaths = new Collection<ExcludedPath>
{
new ExcludedPath { Path = "/contentVector/*" } // Exclude from standard index
}
}
};

await database.CreateContainerIfNotExistsAsync(
containerProperties,
ThroughputProperties.CreateAutoscaleThroughput(4000));
}

Understanding Vector Index Types

Cosmos DB offers three vector index types:

Flat — Small datasets (<10K vectors)— Exact results, higher RU cost
QuantizedFlat — Medium datasets — Good balance of accuracy and cost
DiskANN — Large datasets (millions) — Approximate results, lowest cost

For most applications, start with QuantizedFlat and move to DiskANN as your data grows.

Step 3: Configure Vector Index via JSON (Alternative)

You can also define the indexing policy as JSON:

{
"indexingMode": "consistent",
"automatic": true,
"includedPaths": [
{ "path": "/*" }
],
"excludedPaths": [
{ "path": "/contentVector/*" }
],
"vectorIndexes": [
{
"path": "/contentVector",
"type": "quantizedFlat"
}
]
}

And the vector embedding policy:

{
"vectorEmbeddings": [
{
"path": "/contentVector",
"dataType": "float32",
"dimensions": 1536,
"distanceFunction": "cosine"
}
]
}

Part 2: Generating Embeddings

Before storing notes, we need to generate vector embeddings. Here’s a service using OpenAI:

public interface IEmbeddingService
{
Task<float[]> GetEmbeddingAsync(string text, CancellationToken ct = default);
}

public class OpenAIEmbeddingService : IEmbeddingService
{
private readonly OpenAIClient _client;
private readonly string _model = "text-embedding-3-small"; // 1536 dimensions

public OpenAIEmbeddingService(IConfiguration config)
{
_client = new OpenAIClient(config["OpenAI:ApiKey"]);
}

public async Task<float[]> GetEmbeddingAsync(string text, CancellationToken ct)
{
if (string.IsNullOrWhiteSpace(text))
return Array.Empty<float>();

var embeddingClient = _client.GetEmbeddingClient(_model);
var response = await embeddingClient.GenerateEmbeddingAsync(text, cancellationToken: ct);

return response.Value.ToFloats().ToArray();
}
}

Part 3: Storing Notes with Vectors

When creating a note, generate and store the embedding:

public class NoteService
{
private readonly Container _container;
private readonly IEmbeddingService _embedding;

public async Task<Note> CreateNoteAsync(string content, string partitionKey)
{
// Generate embedding for the content
var vector = await _embedding.GetEmbeddingAsync(content);

var note = new Note
{
Id = $"note-{Guid.NewGuid()}",
PartitionKey = partitionKey,
Content = content,
CreatedAt = DateTime.UtcNow,
ContentVector = vector
};

var response = await _container.CreateItemAsync(note, new PartitionKey(partitionKey));
return response.Resource;
}
}

Part 4: Implementing Search

Full-Text Search

For keyword-based search, use standard Cosmos DB queries:

public async Task<List<Note>> FullTextSearchAsync(
string query,
string partitionKey,
int limit = 20)
{
var sql = @"
SELECT * FROM c
WHERE c.pk = @pk
AND CONTAINS(LOWER(c.content), LOWER(@query))
ORDER BY c.createdAt DESC
OFFSET 0 LIMIT @limit";

var queryDef = new QueryDefinition(sql)
.WithParameter("@pk", partitionKey)
.WithParameter("@query", query)
.WithParameter("@limit", limit);

var results = new List<Note>();
using var iterator = _container.GetItemQueryIterator<Note>(queryDef);

while (iterator.HasMoreResults)
{
var response = await iterator.ReadNextAsync();
results.AddRange(response);
}

return results;
}

Vector (Semantic) Search

Vector search finds semantically similar notes using the “VectorDistance” function:

public async Task<List<NoteWithScore>> VectorSearchAsync(
string query,
string partitionKey,
int limit = 20)
{
// Generate embedding for the search query
var queryVector = await _embedding.GetEmbeddingAsync(query);

// Use VectorDistance for semantic search
// Filter with IS_ARRAY to skip documents without embeddings
var sql = @"
SELECT
c.id,
c.pk,
c.content,
c.createdAt,
VectorDistance(c.contentVector, @queryVector) AS score
FROM c
WHERE c.pk = @pk
AND IS_ARRAY(c.contentVector)
ORDER BY VectorDistance(c.contentVector, @queryVector)
OFFSET 0 LIMIT @limit";

var queryDef = new QueryDefinition(sql)
.WithParameter("@pk", partitionKey)
.WithParameter("@queryVector", queryVector)
.WithParameter("@limit", limit);

var results = new List<NoteWithScore>();
using var iterator = _container.GetItemQueryIterator<NoteWithScore>(queryDef);

while (iterator.HasMoreResults)
{
var response = await iterator.ReadNextAsync();
results.AddRange(response);
}

// Convert distance to similarity (VectorDistance returns lower = more similar)
foreach (var result in results)
{
result.SimilarityScore = 1.0 - result.SimilarityScore;
}

return results;
}

public class NoteWithScore : Note
{
[JsonPropertyName("score")]
public double SimilarityScore { get; set; }
}

How “VectorDistance” Works:

  • Returns the distance between two vectors using the configured distance function (cosine, euclidean, or dot product)
  • For cosine distance: 0 = identical, 2 = completely opposite
  • Lower scores = more similar (convert to similarity with “1.0 — distance”)
  • The query is automatically optimized to use the vector index
  • Use “IS_ARRAY(c.contentVector)” to filter out documents that haven’t been indexed yet

Hybrid Search

Hybrid search combines keyword matching with semantic similarity. This approach works well because:

  • Full-text catches exact matches that users expect
  • Vector search finds conceptually related content
  • The combination reduces both false positives and false negatives
private const int RrfK = 60;  // Standard RRF constant

public async Task<List<NoteWithScore>> HybridSearchAsync(
string query,
string partitionKey,
double textWeight = 0.5,
double vectorWeight = 0.5,
int limit = 20)
{
// Fetch 2x the limit from each search to ensure good coverage for RRF
var expandedLimit = limit * 2;

// Execute both searches in parallel for better performance
var textSearchTask = FullTextSearchAsync(query, partitionKey, expandedLimit);
var vectorSearchTask = VectorSearchAsync(query, partitionKey, expandedLimit);

await Task.WhenAll(textSearchTask, vectorSearchTask);

var textResults = await textSearchTask;
var vectorResults = await vectorSearchTask;

// Apply Reciprocal Rank Fusion (RRF) to combine results
var combinedScores = new Dictionary<string, (Note Note, double Score)>();

// Add text search results with RRF scores
for (int rank = 0; rank < textResults.Count; rank++)
{
var note = textResults[rank];
var rrfScore = textWeight / (RrfK + rank + 1);
combinedScores[note.Id] = (note, rrfScore);
}

// Add vector search results with RRF scores
for (int rank = 0; rank < vectorResults.Count; rank++)
{
var note = vectorResults[rank];
var rrfScore = vectorWeight / (RrfK + rank + 1);

if (combinedScores.TryGetValue(note.Id, out var existing))
{
// Document appears in both result sets - add scores
combinedScores[note.Id] = (existing.Note, existing.Score + rrfScore);
}
else
{
combinedScores[note.Id] = (note, rrfScore);
}
}

// Sort by combined RRF score and take top results
return combinedScores.Values
.OrderByDescending(x => x.Score)
.Take(limit)
.Select(x => new NoteWithScore
{
Id = x.Note.Id,
PartitionKey = x.Note.PartitionKey,
Content = x.Note.Content,
CreatedAt = x.Note.CreatedAt,
SimilarityScore = x.Score
})
.ToList();
}

Why Reciprocal Rank Fusion (RRF)?

RRF is a simple but effective algorithm for combining ranked lists from different sources:

  • The formula “weight / (k + rank + 1)” normalizes scores across different ranking methods
  • The constant “k=60” is a standard value that prevents top results from dominating too heavily
  • Documents appearing in both lists get boosted, which is exactly what we want

Part 5: Putting It All Together

Here’s the complete search service:

public interface ISearchService
{
Task<List<Note>> FullTextSearchAsync(string query, string partitionKey, int limit = 20);
Task<List<NoteWithScore>> VectorSearchAsync(string query, string partitionKey, int limit = 20);
Task<List<NoteWithScore>> HybridSearchAsync(string query, string partitionKey, HybridSearchOptions? options = null);
}

public class HybridSearchOptions
{
public double TextWeight { get; set; } = 0.5;
public double VectorWeight { get; set; } = 0.5;
public int Limit { get; set; } = 20;
}

Register the services in “Program.cs”:

// Cosmos DB
builder.Services.AddSingleton(sp =>
{
var connectionString = builder.Configuration.GetConnectionString("Cosmos");
return new CosmosClient(connectionString, new CosmosClientOptions
{
ApplicationName = "NotesService",
SerializerOptions = new CosmosSerializationOptions
{
PropertyNamingPolicy = CosmosPropertyNamingPolicy.CamelCase
}
});
});

// Embedding service
builder.Services.AddSingleton<IEmbeddingService, OpenAIEmbeddingService>();

// Search service
builder.Services.AddScoped<ISearchService, SearchService>();

Performance Tips

1. Choose the Right Partition Key

Partition your data so searches stay within a single partition when possible. In our example, we use a context-based partition key:

PartitionKey = $"context-{contextId}"  // All notes in a context share a partition

2. Exclude Vectors from Standard Indexing

Vector fields should only be indexed by the vector index, not the standard index:

"excludedPaths": [
{ "path": "/contentVector/*" }
]

3. Use Appropriate Vector Dimensions

Smaller dimensions = faster search and lower storage:

  • “text-embedding-3-small”: 1536 dimensions (good balance)
  • “text-embedding-3-large”: 3072 dimensions (higher accuracy)
  • You can also reduce dimensions via the API

4. Batch Embedding Generation

When indexing many documents, batch your embedding calls:

public async Task<List<float[]>> GetEmbeddingsAsync(List<string> texts, CancellationToken ct)
{
var embeddingClient = _client.GetEmbeddingClient(_model);
var response = await embeddingClient.GenerateEmbeddingsAsync(texts, cancellationToken: ct);

return response.Value
.OrderBy(e => e.Index)
.Select(e => e.ToFloats().ToArray())
.ToList();
}

5. Run Searches in Parallel

When implementing hybrid search, run the text and vector searches concurrently:

// Good: parallel execution
var textTask = FullTextSearchAsync(query, pk, limit);
var vectorTask = VectorSearchAsync(query, pk, limit);
await Task.WhenAll(textTask, vectorTask);

// Bad: sequential execution (2x slower)
var textResults = await FullTextSearchAsync(query, pk, limit);
var vectorResults = await VectorSearchAsync(query, pk, limit);

6. Handle Missing Vectors Gracefully

Not all documents may have embeddings (new documents, indexing failures). Filter them out:

WHERE IS_ARRAY(c.contentVector)  -- Skip documents without vectors

7. Consider Async Indexing

For production systems, index documents asynchronously:

// When creating a note, queue it for indexing
await _indexingQueue.EnqueueAsync(new IndexingJob
{
DocumentId = note.Id,
DocumentType = "Note"
});

// Background service processes the queue
public class IndexingBackgroundService : BackgroundService
{
protected override async Task ExecuteAsync(CancellationToken ct)
{
while (!ct.IsCancellationRequested)
{
var job = await _queue.DequeueAsync(ct);
if (job != null)
{
await _indexer.IndexDocumentAsync(job, ct);
}
}
}
}

Cost Considerations

Vector search operations consume RUs based on:

  • Number of vectors searched
  • Vector dimensions
  • Index type used

Typical costs (approximate):

  • Flat index: ~10–50 RUs per query
  • QuantizedFlat: ~5–20 RUs per query
  • DiskANN: ~2–10 RUs per query

Monitor your RU consumption and adjust throughput accordingly.

Summary

Azure Cosmos DB provides a powerful platform for building intelligent search:

  • Full-text search for exact keyword matching
  • Vector search for semantic similarity using embeddings
  • Hybrid search for the best of both worlds

The key steps are:

  • Enable vector search capability on your Cosmos DB account
  • Configure vector embedding policy and vector indexes on your container
  • Generate embeddings using OpenAI or similar services
  • Store embeddings alongside your documents
  • Use “VectorDistance” for semantic queries

This combination enables search experiences that understand user intent, not just keywords.

References

Read, clap or comment on Medium →