feat(E): Unified Search — 24 Tasks complete
Check Cross-Plugin Imports / check (push) Has been cancelled
Check Cross-Plugin Imports / check (push) Has been cancelled
- SPIKE-E: FTS+Vector+Permission benchmark on 10k records (all <30ms) - E-PROV: supports_fts/vector/rag/graph capability flags on all providers - E-FTS/VEC: All 11 providers refactored to BaseSearchProvider with permission filtering - E-PERM: Over-fetch strategy for vector+permission (15x faster than ANY() filter) - E-FUSE: rrf_fusion_multi() for N-way RRF over FTS+Vector+RAG+Graph - E-LLM: Query understanding cleaned up to use central llm_complete() - E-CHUNK: Document chunking module + document_chunks table with HNSW index - E-EMB: Chunk embedding ARQ jobs (index_file_chunks, reindex_chunks) - E-RAG: RAG retrieval via FileSearchProvider.search_rag() - E-GRAPH: GraphRAG BFS traversal via GraphRAGSearchProvider.search_graph() - E-IX-EVT: Auto-indexing via outbox events + delete/cleanup handlers - E-IX-RE: Batch reindex with progress tracking + reindex_all job - E-DATA-LIFE: Lifecycle module (remove/rebuild/restore/correct) + API endpoints - E-K-MEM: AgentMemorySearchProvider - E-P-AI: AIChatSearchProvider - E-P-WF: WorkflowSearchProvider - E-P-COMM: ConversationSearchProvider verified (already on BaseSearchProvider) - E-API: Filter params (date_from/to, tags, sort) + /facets endpoint - E-TOOL: unified_search AI tool registered in ToolRegistry - E-MCP: Search tool in MCP server with normal RBAC/tenant checks - E-UI-CMD: CommandPalette (Cmd+K) with debounced search + recent searches - E-UI-FAC: SearchFacets, SearchResultCard, SavedSearches components - E-TEST: 40 new tests in test_unified_search_phase_e.py (105 total green) - E-DOC: api-documentation.md, plugin-development-guide.md, test-strategy.md updated 105 tests passing, TypeScript clean.
This commit is contained in:
@@ -25,6 +25,12 @@ class BaseSearchProvider:
|
||||
The base class handles loading visible IDs and passing them to the subclass.
|
||||
"""
|
||||
|
||||
# Capability flags — override in subclass
|
||||
supports_fts: bool = True
|
||||
supports_vector: bool = True
|
||||
supports_rag: bool = False
|
||||
supports_graph: bool = False
|
||||
|
||||
entity_type: str = "" # Override in subclass
|
||||
|
||||
async def search_fts(
|
||||
@@ -54,7 +60,12 @@ class BaseSearchProvider:
|
||||
user_id: uuid.UUID | None = None,
|
||||
is_system_admin: bool = False,
|
||||
) -> list[dict[str, Any]]:
|
||||
"""Semantic vector search with visibility filter."""
|
||||
"""Semantic vector search with visibility filter.
|
||||
|
||||
Uses over-fetch strategy: fetch limit*3 from HNSW without permission
|
||||
filter, then post-filter in Python. This avoids the 15x performance
|
||||
hit of `id = ANY($uuid[])` on HNSW results found in SPIKE-E.
|
||||
"""
|
||||
# Set HNSW ef_search parameter for this transaction
|
||||
await db.execute(text(f"SET LOCAL hnsw.ef_search = {settings.hnsw_ef_search}"))
|
||||
if is_system_admin or not user_id:
|
||||
@@ -63,7 +74,12 @@ class BaseSearchProvider:
|
||||
visible_ids = await self._get_visible_ids(db, tenant_id, user_id)
|
||||
if not visible_ids:
|
||||
return []
|
||||
return await self._search_vector_filtered(db, embedding, tenant_id, limit, visible_ids)
|
||||
|
||||
# Over-fetch 3x the limit, then post-filter in Python
|
||||
over_fetch_limit = limit * 3
|
||||
results = await self._search_vector_filtered(db, embedding, tenant_id, over_fetch_limit, None)
|
||||
filtered = [r for r in results if r.get("id") in visible_ids]
|
||||
return filtered[:limit]
|
||||
|
||||
async def _get_visible_ids(
|
||||
self, db: AsyncSession, tenant_id: uuid.UUID, user_id: uuid.UUID
|
||||
|
||||
Reference in New Issue
Block a user