Search (vector)
Search — pluggable vector store
github.com/redelay/go-modules/search is a backend-agnostic vector-search module. The
public contract — index CRUD, upsert, query — is identical across backends; the first
shipping backend is Qdrant, with OpenSearch planned. Embeddings are computed via any
registered go-ai LLM provider so callers can switch between OVH bge-m3 (€0.01/M),
OpenAI text-embedding-3-small, or a local Ollama model without touching application
code.
go get github.com/redelay/go-modules/search
Blank-import the module + a backend + the FlowDSL node pack:
import (
_ "github.com/redelay/go-modules/search" // framework module + HTTP surface
_ "github.com/redelay/go-modules/search/qdrant" // Qdrant backend factory
_ "github.com/redelay/go-modules/search/flowdsl" // FlowDSL nodes: search-upsert, search-query
)
The admin submodule at search/admin mounts index CRUD under the admin prefix — import
it from cmd/admin-api only.
Concepts
| Term | Meaning |
|---|---|
| Index | A named collection of vectors. Callers use a short logical name (docs); the service rewrites it to {SEARCH_INDEX_PREFIX}{name} before talking to the backend. |
| Document | One record — {id, text?, vector?, payload?}. Exactly one of text or vector must be present; the service auto-embeds text via the configured LLM provider. |
| Query | Nearest-neighbour request — {text?, vector?, limit?, payload?, filter?}. text is auto-embedded. payload is a shortcut for exact-match filters; filter is forwarded to the backend verbatim for richer expressions (Qdrant Filter, OpenSearch knn_filter, …). |
| Hit | One result — {id, score, payload?}. |
Index prefix — safe multi-tenancy
SEARCH_INDEX_PREFIX is stamped onto every index name before it reaches the backend, and
stripped off on the way back out. One Qdrant cluster can host redelay_dev_docs,
redelay_prod_docs, tenant_a_docs simultaneously, with each app seeing only its own
indexes via ListIndexes. Change the prefix at runtime via the admin setting
search.index_prefix — existing indexes under the old prefix become invisible until
moved.
Configuration
Config (ENV, startup):
| Key | Default | Purpose |
|---|---|---|
SEARCH_BACKEND | qdrant | Backend factory id — qdrant, later opensearch |
SEARCH_INDEX_PREFIX | redelay_ | Stamped onto every index name |
SEARCH_EMBEDDING_PROVIDER | ollama | Registered llm provider id used for auto-embedding |
SEARCH_EMBEDDING_MODEL | nomic-embed-text | Embedding model name |
SEARCH_EMBEDDING_DIM | 768 | Dimension — must match the model (bge-m3 = 1024, ada-002 = 1536, nomic-embed = 768) |
QDRANT_URL | http://localhost:6333 | Qdrant HTTP endpoint |
QDRANT_API_KEY | (empty) | Optional — sent as api-key header |
Admin settings (runtime, override env):
| Key | Default | Purpose |
|---|---|---|
search.default_distance | cosine | Similarity metric for new indexes |
search.default_limit | 10 | Top-K when callers omit limit |
search.index_prefix | (uses env) | Runtime override |
search.embedding_provider | (uses env) | Runtime override |
search.embedding_model | (uses env) | Runtime override |
Settings appear in /admin/settings/search — the same module pattern as ai-guard and
scheduler.
HTTP surface
Public (/search/* — mounted on both binaries):
| Method | Path | Purpose |
|---|---|---|
| GET | /search/status | Module configuration snapshot (backend id, prefix, embedding provider/model/dim) |
| GET | /search/indexes | List caller-visible indexes (prefix-stripped) |
| POST | /search/indexes/{name}/query | Vector query by text or vector |
Admin (/api/v1/admin/search/* — admin-api only):
| Method | Path | Purpose |
|---|---|---|
| GET | /admin/search/indexes | List |
| POST | /admin/search/indexes | Create (body: {name, dim, distance?}) |
| GET | /admin/search/indexes/{name} | Metadata — dimension, distance, points count |
| DELETE | /admin/search/indexes/{name} | Remove index + all vectors |
| POST | /admin/search/indexes/{name}/upsert | Upsert documents — body {documents: [...]} |
| POST | /admin/search/indexes/{name}/delete | Delete by id — body {ids: [...]} |
| POST | /admin/search/indexes/{name}/reindex | Rebuild the index from its source of truth (see below) |
Permissions defined: search:index:list, search:index:write, search:doc:write,
search:doc:delete.
Keeping an index fresh
An index is only as good as its last write. Two mechanisms keep it current, and a well-behaved owner module uses both.
1. Reindex on change (incremental, preferred)
Have the owning collection publish lifecycle events and consume them. Any store
built on crud.MongoCRUD opts in with one call:
// programs/crud.go — emits program.created / .updated / .deleted
crud.NewMongoCRUD[*WorkoutProgram](db.Collection(coll)).WithEvents(bus, "program")
Then react to exactly the entity that changed:
func (m *Module) Events() []*ir.Event { return nil } // required alongside Consumers()
func (m *Module) Consumers() []*modules.ConsumerRegistration {
return []*modules.ConsumerRegistration{
{Topic: "program.updated", GroupID: "my-reindex", Handler: m.onChange},
}
}
func (m *Module) onChange(ctx context.Context, ev *modules.EventMessage) error {
doc, ok := m.buildDoc(ctx, ev.EntityID) // still published?
if !ok {
return search.Current().Delete(ctx, "programs", []string{ev.EntityID})
}
return search.Current().Upsert(ctx, "programs", []search.Document{doc})
}
This makes a newly created, edited, translated, or unpublished entity searchable (or gone) within a second — no polling. Deleting the point when the entity is no longer public is what stops unpublished rows leaking into results.
Gotcha:
modules.EventsProviderrequires bothEvents()andConsumers(). Implement onlyConsumers()and the type assertion fails silently — no consumer is ever registered.
2. Reindex on demand (full rebuild)
Register a rebuild callback so operators get a Reindex action in the admin — needed after a bulk translation, a schema change, or switching embedding model (a new model means new vectors for every document):
// Startup
search.RegisterReindexer("programs", m) // m implements search.Reindexer
func (m *Module) Reindex(ctx context.Context, index string) (int, error) {
// re-embed + upsert everything; return the document count
}
GET /admin/search/entities then lists the index in reindexable, and
POST /admin/search/indexes/programs/reindex returns
{"index":"programs","documents":5,"status":"ok"}. Indexes without a registered
owner return 501 Not Implemented rather than silently doing nothing.
A periodic full reindex (a ticker in the owner module) is still a cheap safety net against missed events — embeddings cost fractions of a cent per thousand documents.
Prune what no longer exists
A rebuild writes what currently exists and knows nothing about what was
deleted, so upserting alone lets the index drift: an unpublished or deleted
row keeps answering searches forever. Close the loop with PruneStale:
ids := make([]string, 0, len(docs))
for _, d := range docs { ids = append(ids, d.ID) }
svc.Upsert(ctx, index, docs)
svc.PruneStale(ctx, index, ids) // removes every point not in ids
It pages through the index, diffs against the ids you just wrote, and deletes the remainder — returning how many it removed. Treat it as best-effort: a prune failure should not fail a rebuild that already succeeded.
Without this, deleting rows straight in the database (a cleanup script, a migration) leaves orphans that no event will ever reap.
Backend — Qdrant
The bundled backend talks to Qdrant over its REST API — no SDK, no protobuf dep.
- Collections map 1:1 with indexes.
CreateIndexis idempotent on matching params; disagreeing dimensions return a 400 from Qdrant. - Point IDs — Qdrant accepts only unsigned integers or UUIDs. The backend transparently
hashes arbitrary string IDs (e.g.
docs/admin/overview.md) to deterministic UUIDv5 values and stashes the original underpayload._id. Search results restore the original string so callers see stable IDs regardless of backend encoding. Decimal-string IDs pass through as numbers. - Filtering —
Query.Payloadis a shortcut formustequality constraints;Query.Filteris forwarded verbatim for richer expressions.
FlowDSL nodes (search/flowdsl)
| Node ID | Kind | Purpose |
|---|---|---|
redelay/search-upsert | action | Index one or more documents. Empty vectors are auto-embedded. |
redelay/search-query | action | Semantic query by text (auto-embedded) or vector. |
Both nodes resolve the live *search.Service lazily via search.Current() so handler
registration can happen before the module's Startup. Handlers are wired in
backend/imports/aiwire/aiwire.go alongside the llm-chat / llm-embed / llm-guard
handlers.
Example flow — RAG over docs:
# upsert pipeline
- id: chunk
action_ref: redelay/text-split # project node
- id: index
action_ref: redelay/search-upsert
config: { index: docs }
# query pipeline
- id: search
action_ref: redelay/search-query
config: { index: docs, limit: 5 }
- id: build-prompt
action_ref: redelay/template-format # fold hits into a system prompt
- id: answer
action_ref: redelay/llm-chat
Integration testing
go-modules/search/qdrant/integration_test.go runs end-to-end against a live Qdrant —
create collection, upsert three points, nearest-neighbour query, verify the closest match,
tear down. Gated on QDRANT_URL:
QDRANT_URL=http://100.93.236.60:6333 go test -run Integration -v ./search/qdrant/
Pair with OVH_AI_ENDPOINTS_TOKEN set to exercise the full pipeline — OVH bge-m3
embedding → Qdrant upsert → OVH chat retrieval.
Coordination
Cross-container coordination layer — pub/sub invalidation, fan-out streams, and watchable KV for multi-container Redelay deployments.
Profiles
Reusable, admin-editable config bundles — one dropdown on a FlowDSL node replaces eight hand-configured fields. Works for LLM chat, guard, databases, FTP, storage, any connection-bearing node.