Reference

Search (vector)

Pluggable vector-search module — Qdrant today, OpenSearch planned. Index CRUD, semantic query, embedding-provider switching, FlowDSL nodes.

Search — pluggable vector store

github.com/redelay/go-modules/search is a backend-agnostic vector-search module. The public contract — index CRUD, upsert, query — is identical across backends; the first shipping backend is Qdrant, with OpenSearch planned. Embeddings are computed via any registered go-ai LLM provider so callers can switch between OVH bge-m3 (€0.01/M), OpenAI text-embedding-3-small, or a local Ollama model without touching application code.

shell
go get github.com/redelay/go-modules/search

Blank-import the module + a backend + the FlowDSL node pack:

go
import (
    _ "github.com/redelay/go-modules/search"          // framework module + HTTP surface
    _ "github.com/redelay/go-modules/search/qdrant"   // Qdrant backend factory
    _ "github.com/redelay/go-modules/search/flowdsl"  // FlowDSL nodes: search-upsert, search-query
)

The admin submodule at search/admin mounts index CRUD under the admin prefix — import it from cmd/admin-api only.

Concepts

TermMeaning
IndexA named collection of vectors. Callers use a short logical name (docs); the service rewrites it to {SEARCH_INDEX_PREFIX}{name} before talking to the backend.
DocumentOne record — {id, text?, vector?, payload?}. Exactly one of text or vector must be present; the service auto-embeds text via the configured LLM provider.
QueryNearest-neighbour request — {text?, vector?, limit?, payload?, filter?}. text is auto-embedded. payload is a shortcut for exact-match filters; filter is forwarded to the backend verbatim for richer expressions (Qdrant Filter, OpenSearch knn_filter, …).
HitOne result — {id, score, payload?}.

Index prefix — safe multi-tenancy

SEARCH_INDEX_PREFIX is stamped onto every index name before it reaches the backend, and stripped off on the way back out. One Qdrant cluster can host redelay_dev_docs, redelay_prod_docs, tenant_a_docs simultaneously, with each app seeing only its own indexes via ListIndexes. Change the prefix at runtime via the admin setting search.index_prefix — existing indexes under the old prefix become invisible until moved.

Configuration

Config (ENV, startup):

KeyDefaultPurpose
SEARCH_BACKENDqdrantBackend factory id — qdrant, later opensearch
SEARCH_INDEX_PREFIXredelay_Stamped onto every index name
SEARCH_EMBEDDING_PROVIDERollamaRegistered llm provider id used for auto-embedding
SEARCH_EMBEDDING_MODELnomic-embed-textEmbedding model name
SEARCH_EMBEDDING_DIM768Dimension — must match the model (bge-m3 = 1024, ada-002 = 1536, nomic-embed = 768)
QDRANT_URLhttp://localhost:6333Qdrant HTTP endpoint
QDRANT_API_KEY(empty)Optional — sent as api-key header

Admin settings (runtime, override env):

KeyDefaultPurpose
search.default_distancecosineSimilarity metric for new indexes
search.default_limit10Top-K when callers omit limit
search.index_prefix(uses env)Runtime override
search.embedding_provider(uses env)Runtime override
search.embedding_model(uses env)Runtime override

Settings appear in /admin/settings/search — the same module pattern as ai-guard and scheduler.

HTTP surface

Public (/search/* — mounted on both binaries):

MethodPathPurpose
GET/search/statusModule configuration snapshot (backend id, prefix, embedding provider/model/dim)
GET/search/indexesList caller-visible indexes (prefix-stripped)
POST/search/indexes/{name}/queryVector query by text or vector

Admin (/api/v1/admin/search/* — admin-api only):

MethodPathPurpose
GET/admin/search/indexesList
POST/admin/search/indexesCreate (body: {name, dim, distance?})
GET/admin/search/indexes/{name}Metadata — dimension, distance, points count
DELETE/admin/search/indexes/{name}Remove index + all vectors
POST/admin/search/indexes/{name}/upsertUpsert documents — body {documents: [...]}
POST/admin/search/indexes/{name}/deleteDelete by id — body {ids: [...]}
POST/admin/search/indexes/{name}/reindexRebuild the index from its source of truth (see below)

Permissions defined: search:index:list, search:index:write, search:doc:write, search:doc:delete.

Keeping an index fresh

An index is only as good as its last write. Two mechanisms keep it current, and a well-behaved owner module uses both.

1. Reindex on change (incremental, preferred)

Have the owning collection publish lifecycle events and consume them. Any store built on crud.MongoCRUD opts in with one call:

go
// programs/crud.go — emits program.created / .updated / .deleted
crud.NewMongoCRUD[*WorkoutProgram](db.Collection(coll)).WithEvents(bus, "program")

Then react to exactly the entity that changed:

go
func (m *Module) Events() []*ir.Event { return nil } // required alongside Consumers()

func (m *Module) Consumers() []*modules.ConsumerRegistration {
    return []*modules.ConsumerRegistration{
        {Topic: "program.updated", GroupID: "my-reindex", Handler: m.onChange},
    }
}

func (m *Module) onChange(ctx context.Context, ev *modules.EventMessage) error {
    doc, ok := m.buildDoc(ctx, ev.EntityID) // still published?
    if !ok {
        return search.Current().Delete(ctx, "programs", []string{ev.EntityID})
    }
    return search.Current().Upsert(ctx, "programs", []search.Document{doc})
}

This makes a newly created, edited, translated, or unpublished entity searchable (or gone) within a second — no polling. Deleting the point when the entity is no longer public is what stops unpublished rows leaking into results.

Gotcha: modules.EventsProvider requires both Events() and Consumers(). Implement only Consumers() and the type assertion fails silently — no consumer is ever registered.

2. Reindex on demand (full rebuild)

Register a rebuild callback so operators get a Reindex action in the admin — needed after a bulk translation, a schema change, or switching embedding model (a new model means new vectors for every document):

go
// Startup
search.RegisterReindexer("programs", m) // m implements search.Reindexer

func (m *Module) Reindex(ctx context.Context, index string) (int, error) {
    // re-embed + upsert everything; return the document count
}

GET /admin/search/entities then lists the index in reindexable, and POST /admin/search/indexes/programs/reindex returns {"index":"programs","documents":5,"status":"ok"}. Indexes without a registered owner return 501 Not Implemented rather than silently doing nothing.

A periodic full reindex (a ticker in the owner module) is still a cheap safety net against missed events — embeddings cost fractions of a cent per thousand documents.

Prune what no longer exists

A rebuild writes what currently exists and knows nothing about what was deleted, so upserting alone lets the index drift: an unpublished or deleted row keeps answering searches forever. Close the loop with PruneStale:

go
ids := make([]string, 0, len(docs))
for _, d := range docs { ids = append(ids, d.ID) }
svc.Upsert(ctx, index, docs)
svc.PruneStale(ctx, index, ids)   // removes every point not in ids

It pages through the index, diffs against the ids you just wrote, and deletes the remainder — returning how many it removed. Treat it as best-effort: a prune failure should not fail a rebuild that already succeeded.

Without this, deleting rows straight in the database (a cleanup script, a migration) leaves orphans that no event will ever reap.

Backend — Qdrant

The bundled backend talks to Qdrant over its REST API — no SDK, no protobuf dep.

  • Collections map 1:1 with indexes. CreateIndex is idempotent on matching params; disagreeing dimensions return a 400 from Qdrant.
  • Point IDs — Qdrant accepts only unsigned integers or UUIDs. The backend transparently hashes arbitrary string IDs (e.g. docs/admin/overview.md) to deterministic UUIDv5 values and stashes the original under payload._id. Search results restore the original string so callers see stable IDs regardless of backend encoding. Decimal-string IDs pass through as numbers.
  • Filtering — Query.Payload is a shortcut for must equality constraints; Query.Filter is forwarded verbatim for richer expressions.

FlowDSL nodes (search/flowdsl)

Node IDKindPurpose
redelay/search-upsertactionIndex one or more documents. Empty vectors are auto-embedded.
redelay/search-queryactionSemantic query by text (auto-embedded) or vector.

Both nodes resolve the live *search.Service lazily via search.Current() so handler registration can happen before the module's Startup. Handlers are wired in backend/imports/aiwire/aiwire.go alongside the llm-chat / llm-embed / llm-guard handlers.

Example flow — RAG over docs:

yaml
# upsert pipeline
- id: chunk
  action_ref: redelay/text-split    # project node
- id: index
  action_ref: redelay/search-upsert
  config: { index: docs }

# query pipeline
- id: search
  action_ref: redelay/search-query
  config: { index: docs, limit: 5 }
- id: build-prompt
  action_ref: redelay/template-format   # fold hits into a system prompt
- id: answer
  action_ref: redelay/llm-chat

Integration testing

go-modules/search/qdrant/integration_test.go runs end-to-end against a live Qdrant — create collection, upsert three points, nearest-neighbour query, verify the closest match, tear down. Gated on QDRANT_URL:

shell
QDRANT_URL=http://100.93.236.60:6333 go test -run Integration -v ./search/qdrant/

Pair with OVH_AI_ENDPOINTS_TOKEN set to exercise the full pipeline — OVH bge-m3 embedding → Qdrant upsert → OVH chat retrieval.