feat(B-SENS): Sensitive Data Boundary + AI/Data Exposure Policy + AIProvider Compliance
Check Cross-Plugin Imports / check (push) Has been cancelled
Check Cross-Plugin Imports / check (push) Has been cancelled
B-SENS: app/core/sensitive_data.py (NEU) — zentrale Sensitive-Field-Verwaltung - SENSITIVE_FIELDS dict für contact/user/mail_account/system_settings - is_sensitive(), sanitize_dict(), register_sensitive_fields() - Integration: errors.py (Log-Redaction), audit.py (Audit-Masking), export_service.py (Export-Filter), embedding.py (Index-Filter) B-DATA-POL: AI/Data Exposure Policy - DATA_EXPOSURE_POLICY: pro Entity+Field welche Systeme erlaubt (llm_context/search/embeddings/rag/agent_memory/export) - filter_for_llm_context/search/embeddings/export/rag/agent_memory() B-AIPROV-COMP: AIProvider Compliance Metadata - Migration 0119: 7 neue Spalten an ai_providers (region, hosting_type, dpa_status, retention_policy, training_on_customer_data, transfer_notice, allowed_data_classes) - llm_client.py: get_provider_compliance() + check_data_class_allowed() B-PRIV-TEST: 76 Tests in test_sensitive_data.py — alle grün - Sensitive Fields, Exposure Policy, Provider Compliance, Secrets-always-blocked - Keine Regression: 39 LLM-Client Tests grün
This commit is contained in:
@@ -147,6 +147,29 @@ async def index_entity(
|
||||
logger.debug("Empty embedding text for %s/%s", entity_type, entity_id)
|
||||
return False
|
||||
|
||||
# Apply sensitive-data filter: ensure no sensitive fields leak into
|
||||
# embedding text. The provider builds text from DB columns, so we
|
||||
# rely on the provider selecting only non-sensitive columns. This
|
||||
# is a secondary safety net — providers should use filter_for_embeddings
|
||||
# when constructing embedding text from dict-like data.
|
||||
from app.core.sensitive_data import get_sensitive_fields
|
||||
|
||||
sensitive = get_sensitive_fields(entity_type)
|
||||
if sensitive:
|
||||
# If any sensitive field name appears as a substring in the text,
|
||||
# it's likely a key=value pair — redact it. This is a best-effort
|
||||
# guard; providers are expected to exclude sensitive columns at
|
||||
# the SQL level.
|
||||
for field_name in sensitive:
|
||||
# Only redact if the field name appears as a key-like pattern
|
||||
import re
|
||||
text = re.sub(
|
||||
rf"\b{re.escape(field_name)}\s*[=:]\s*\S+",
|
||||
f"{field_name}=***REDACTED***",
|
||||
text,
|
||||
flags=re.IGNORECASE,
|
||||
)
|
||||
|
||||
embedding = await generate_embedding(text, db=db, tenant_id=tenant_id)
|
||||
if not embedding:
|
||||
return False
|
||||
|
||||
Reference in New Issue
Block a user