Table of Contents

Backend Setup

In-Memory

The simplest backend — no external dependencies. Data is held in RAM and lost when the process exits. Good for development, tests, and demos.

dotnet add package Mythosia.VectorDb.InMemory
using Mythosia.VectorDb;
using Mythosia.VectorDb.InMemory;

var store = new InMemoryVectorStore();

Built-in hybrid search: RRF (Reciprocal Rank Fusion) merges cosine similarity and BM25 keyword scores.

Concurrent use and record ownership

When a shared store is updated during a query, the stored text and keyword index must describe the same revision. InMemoryVectorStore synchronizes writes, deletes and reads so each vector, text or hybrid query sees a consistent state. Both parts of a hybrid query use that same state.

The store copies incoming records, including vector arrays and metadata, and returns independent copies from lookups, searches and diagnostics. Editing an input or returned record does not change stored data; call UpsertAsync again to save the change. Keep input records, vectors and metadata unchanged until the call finishes copying or reading them.

A supplied CancellationToken can cancel a call while it waits for another operation to release the store lock. Canceling that wait does not itself abort the operation currently using the store. Cancellation does not split a record’s body/index update once that update has begun.

A canceled batch can retain records already written. ReplaceByFilterAsync still performs deletion followed by batch insertion without a transaction: another query can observe the gap, and failures or cancellation do not roll back earlier writes.

Diagnostics

IVectorStoreDiagnostics is an optional contract in Mythosia.VectorDb.Abstractions 4.2.0. InMemory implements it directly; IVectorStore gains no required members. ListAllRecordsAsync lists all records and ScoredListAsync returns all similarity scores in descending order without TopK filtering. Both inspect the whole store: they accept no metadata filter and do not apply a RAG pipeline’s StoreFilter. GetTotalRecordCount() remains an InMemory convenience method outside the contract. RAG-specific chunk analysis, health checks and reports remain in RagDiagnostics and RagDiagnosticSession in Mythosia.AI.Rag.

This release intentionally includes a breaking interface migration in the minor versions RAG 8.3.0 and InMemory 4.3.0. Treat it as a release-specific versioning exception: existing InMemory callers using IRagDiagnosticsStore must migrate to IVectorStoreDiagnostics even though the major numbers remain unchanged.

Upgrading to InMemory 4.3.0: use RAG 8.3.0 alongside it; older RAG packages with the new InMemory package are unsupported. InMemory no longer implements IRagDiagnosticsStore: migrate assignments, casts and capability checks to IVectorStoreDiagnostics, and rebuild affected consumers. RAG Abstractions 6.5.0 keeps the obsolete interface, its two original method declarations and default bridges for legacy custom implementations. That bridge does not restore InMemory’s old interface relationship or guarantee compatibility for every old binary.

For custom stores implementing IRagDiagnosticsStore, RAG uses an internal adapter to call the original interface methods. This preserves explicit implementations even when public helper methods have the same signatures. Calling methods through a direct cast to IVectorStoreDiagnostics can select those public methods instead.

// List all stored records
IVectorStoreDiagnostics diagnostics = store;
var all = await diagnostics.ListAllRecordsAsync();
Console.WriteLine($"Total: {store.GetTotalRecordCount()}");

// Inspect raw similarity scores
var scored = await diagnostics.ScoredListAsync(queryVector);
foreach (var r in scored)
    Console.WriteLine($"[{r.Score:F3}] {r.Record.Content[..60]}");

Qdrant

Production-grade vector database with native hybrid search. Runs as a standalone service via Docker or Qdrant Cloud.

dotnet add package Mythosia.VectorDb.Qdrant
# Start Qdrant locally
docker run -p 6333:6333 -p 6334:6334 qdrant/qdrant
using Mythosia.VectorDb.Qdrant;

var store = new QdrantStore(new QdrantOptions
{
    Host             = "localhost",
    Port             = 6334,           // gRPC port
    CollectionName   = "my-docs",
    Dimension        = 1536,           // Must match your embedding model
    AutoCreateCollection = true        // Creates collection on first upsert
});

All Options

new QdrantOptions
{
    Host                   = "localhost",
    Port                   = 6334,
    UseTls                 = false,
    ApiKey                 = null,             // Required for Qdrant Cloud

    CollectionName         = "my-collection",  // Required
    Dimension              = 1536,             // Required

    DistanceStrategy       = QdrantDistanceStrategy.Cosine,
    HybridFusionStrategy   = QdrantHybridFusionStrategy.Rrf,
    AutoCreateCollection   = true,

    // Add extra payload indexes for faster server-side filtering
    AdditionalPayloadIndexes = new List<QdrantIndexOption>
    {
        new QdrantIndexOption { Field = "meta.language", SchemaType = PayloadSchemaType.Keyword },
        new QdrantIndexOption { Field = "meta.date",     SchemaType = PayloadSchemaType.Integer }
    }
}

Distance Strategies

Value Description
Cosine Cosine similarity — best for normalized embeddings (default)
Euclidean L2 distance — lower distance = more similar
DotProduct Dot product — use with unit-normalized vectors

Hybrid Fusion Strategies

Value Description
Rrf Reciprocal Rank Fusion — robust rank-based merging (default)
Dbsf Distribution-Based Score Fusion — merges by score distribution

Qdrant Cloud

new QdrantOptions
{
    Host           = "your-cluster.cloud.qdrant.io",
    Port           = 6334,
    UseTls         = true,
    ApiKey         = "your-qdrant-cloud-key",
    CollectionName = "production",
    Dimension      = 1536
}

Using an External QdrantClient

If you already have a configured QdrantClient (e.g., from a DI container), pass it directly:

var store = new QdrantStore(options, existingQdrantClient);

The store will not dispose the externally provided client.

All vector stores implement IDisposable. When you create a store with the standard constructor, call Dispose() (or use using) to release internal resources.


Pinecone

Fully managed serverless vector database. No infrastructure to manage.

dotnet add package Mythosia.VectorDb.Pinecone
using Mythosia.VectorDb.Pinecone;

var store = new PineconeStore(new PineconeOptions
{
    IndexHost = "https://my-index-xxxx.svc.us-east1-gcp.pinecone.io",
    ApiKey    = "your-api-key"
});

Auto-Create Index

If you don't have an index yet, let the SDK create it:

new PineconeOptions
{
    ApiKey          = "your-api-key",
    AutoCreateIndex = true,
    IndexName       = "my-index",
    Dimension       = 1536,
    Cloud           = "aws",          // "aws", "gcp", or "azure"
    Region          = "us-east-1"
}

When AutoCreateIndex is enabled, the index is created with dotproduct metric — required for hybrid (sparse + dense) search.

All Options

new PineconeOptions
{
    IndexHost              = "https://...",   // Required (or use AutoCreateIndex)
    ApiKey                 = "...",           // Required
    Namespace              = "production",    // Optional: applied to all operations

    UpsertBatchSize        = 100,             // Records per batch upsert request
    RequestTimeoutSeconds  = 100,

    AutoCreateIndex        = false,
    IndexName              = null,
    Dimension              = 0,
    Cloud                  = null,
    Region                 = null,
    ControlPlaneHost       = "https://api.pinecone.io"
}

Using an External HttpClient

If you already have a configured HttpClient (e.g., from IHttpClientFactory):

var store = new PineconeStore(options, existingHttpClient);

The store will not dispose the externally provided client.


PostgreSQL (pgvector)

Uses the pgvector extension to add vector similarity search to a standard PostgreSQL database.

dotnet add package Mythosia.VectorDb.Postgres

Prerequisites

-- Run once on your PostgreSQL server
CREATE EXTENSION IF NOT EXISTS vector;
CREATE EXTENSION IF NOT EXISTS pg_trgm;  -- Only if using Trigram text search

Or let the SDK handle it automatically with EnsureSchema = true.

using Mythosia.VectorDb.Postgres;

var store = new PostgresStore(new PostgresOptions
{
    ConnectionString = "Host=localhost;Port=5432;Database=mydb;Username=user;Password=pass;",
    Dimension        = 1536,
    EnsureSchema     = true    // Auto-creates extension, table, and indexes
});

Index Types

Type Class When to Use
HNSW HnswIndexOptions Default. Fast approximate search. Best for most use cases.
IVFFlat IvfFlatIndexOptions Lower memory. Good for large static datasets.
None NoIndexOptions Sequential scan. Use only for tiny datasets.
// HNSW (default)
new PostgresOptions
{
    // ...
    Index = new HnswIndexOptions
    {
        M              = 16,   // Max neighbor connections per node
        EfConstruction = 64,   // Search scope during index build (higher = better quality)
        EfSearch       = 40    // Runtime search scope (higher = better recall, slower)
    }
}

// IVFFlat
new PostgresOptions
{
    // ...
    Index = new IvfFlatIndexOptions
    {
        Lists  = 100,  // Number of inverted lists
        Probes = 10    // How many lists to probe at query time
    }
}

// No index (sequential scan)
new PostgresOptions { Index = new NoIndexOptions() }

Text Search Modes

Used for the keyword side of hybrid search:

Mode Best For
TsVector Standard full-text search — English, most Western languages
Trigram CJK languages (Korean, Chinese, Japanese), fuzzy matching
new PostgresOptions
{
    TextSearchMode   = TextSearchMode.Trigram,
    TextSearchConfig = "simple"     // PostgreSQL text search configuration
}

Distance Strategies

Value Postgres Operator Notes
Cosine <=> 1 − cosine similarity (default)
Euclidean <-> L2 distance
InnerProduct <#> Negative dot product — use with unit-normalized vectors

Runtime Search Profile

Fine-tune recall vs. latency at query time:

var opts = new HnswSearchRuntimeOptions
{
    Profile = SearchProfile.HighRecall,  // Fast | Balanced | HighRecall
    EfSearch = 80                        // Override HNSW ef_search directly
};

var results = await store.SearchAsync(queryVector, topK: 5, filter: null, runtimeOptions: opts);

Requires Mythosia.VectorDb.Postgres 10.8.1 or later for configured vector-search settings to apply when both hybrid legs are active. Patch notes.

To control the vector candidate search in hybrid retrieval, set HnswIndexOptions.EfSearch or IvfFlatIndexOptions.Probes on PostgresOptions.Index. These defaults apply to ordinary vector search and the vector leg of hybrid search within the search transaction. The per-request runtime overrides shown above apply only to SearchAsync. Approximate search and filtering can still return fewer than topK results.

All Options

new PostgresOptions
{
    ConnectionString  = "...",
    Dimension         = 1536,

    SchemaName        = "public",
    TableName         = "vectors",

    EnsureSchema      = false,
    DistanceStrategy  = DistanceStrategy.Cosine,
    Index             = new HnswIndexOptions(),

    TextSearchConfig  = "simple",
    TextSearchMode    = TextSearchMode.TsVector,

    FailFastOnIndexCreationFailure = true
}