feat(B-SENS): Sensitive Data Boundary + AI/Data Exposure Policy + AIProvider Compliance
Check Cross-Plugin Imports / check (push) Has been cancelled

B-SENS: app/core/sensitive_data.py (NEU) — zentrale Sensitive-Field-Verwaltung
- SENSITIVE_FIELDS dict für contact/user/mail_account/system_settings
- is_sensitive(), sanitize_dict(), register_sensitive_fields()
- Integration: errors.py (Log-Redaction), audit.py (Audit-Masking), export_service.py (Export-Filter), embedding.py (Index-Filter)

B-DATA-POL: AI/Data Exposure Policy
- DATA_EXPOSURE_POLICY: pro Entity+Field welche Systeme erlaubt (llm_context/search/embeddings/rag/agent_memory/export)
- filter_for_llm_context/search/embeddings/export/rag/agent_memory()

B-AIPROV-COMP: AIProvider Compliance Metadata
- Migration 0119: 7 neue Spalten an ai_providers (region, hosting_type, dpa_status, retention_policy, training_on_customer_data, transfer_notice, allowed_data_classes)
- llm_client.py: get_provider_compliance() + check_data_class_allowed()

B-PRIV-TEST: 76 Tests in test_sensitive_data.py — alle grün
- Sensitive Fields, Exposure Policy, Provider Compliance, Secrets-always-blocked
- Keine Regression: 39 LLM-Client Tests grün
This commit is contained in:
Agent Zero
2026-08-13 20:39:32 +02:00
parent bb36378494
commit b231c2d0d3
10 changed files with 939 additions and 15 deletions
@@ -43,6 +43,15 @@ class AIProvider(Base, TenantMixin):
is_default: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
config: Mapped[dict] = mapped_column(JSONB, nullable=False, default=dict)
# ── Compliance metadata (B-AIPROV-COMP) ──
region: Mapped[str] = mapped_column(String(20), nullable=False, default="unknown")
hosting_type: Mapped[str] = mapped_column(String(30), nullable=False, default="cloud")
dpa_status: Mapped[str] = mapped_column(String(20), nullable=False, default="none")
retention_policy: Mapped[str] = mapped_column(Text, nullable=False, default="")
training_on_customer_data: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
transfer_notice: Mapped[str] = mapped_column(Text, nullable=False, default="")
allowed_data_classes: Mapped[list] = mapped_column(JSONB, nullable=False, default=list)
# --- Models ---
@@ -18,6 +18,14 @@ class AIProviderCreate(BaseModel):
is_active: bool = True
is_default: bool = False
config: dict[str, Any] = Field(default_factory=dict)
# Compliance metadata (B-AIPROV-COMP)
region: str = Field(default="unknown", max_length=20)
hosting_type: str = Field(default="cloud", max_length=30)
dpa_status: str = Field(default="none", max_length=20)
retention_policy: str = Field(default="")
training_on_customer_data: bool = False
transfer_notice: str = Field(default="")
allowed_data_classes: list[str] = Field(default_factory=list)
class AIProviderUpdate(BaseModel):
@@ -28,6 +36,14 @@ class AIProviderUpdate(BaseModel):
is_active: bool | None = None
is_default: bool | None = None
config: dict[str, Any] | None = None
# Compliance metadata (B-AIPROV-COMP)
region: str | None = Field(None, max_length=20)
hosting_type: str | None = Field(None, max_length=30)
dpa_status: str | None = Field(None, max_length=20)
retention_policy: str | None = None
training_on_customer_data: bool | None = None
transfer_notice: str | None = None
allowed_data_classes: list[str] | None = None
class AIProviderResponse(BaseModel):
@@ -40,6 +56,14 @@ class AIProviderResponse(BaseModel):
is_active: bool
is_default: bool
config: dict[str, Any]
# Compliance metadata (B-AIPROV-COMP)
region: str = "unknown"
hosting_type: str = "cloud"
dpa_status: str = "none"
retention_policy: str = ""
training_on_customer_data: bool = False
transfer_notice: str = ""
allowed_data_classes: list[str] = Field(default_factory=list)
created_at: datetime | None = None
updated_at: datetime | None = None
@@ -147,6 +147,29 @@ async def index_entity(
logger.debug("Empty embedding text for %s/%s", entity_type, entity_id)
return False
# Apply sensitive-data filter: ensure no sensitive fields leak into
# embedding text. The provider builds text from DB columns, so we
# rely on the provider selecting only non-sensitive columns. This
# is a secondary safety net — providers should use filter_for_embeddings
# when constructing embedding text from dict-like data.
from app.core.sensitive_data import get_sensitive_fields
sensitive = get_sensitive_fields(entity_type)
if sensitive:
# If any sensitive field name appears as a substring in the text,
# it's likely a key=value pair — redact it. This is a best-effort
# guard; providers are expected to exclude sensitive columns at
# the SQL level.
for field_name in sensitive:
# Only redact if the field name appears as a key-like pattern
import re
text = re.sub(
rf"\b{re.escape(field_name)}\s*[=:]\s*\S+",
f"{field_name}=***REDACTED***",
text,
flags=re.IGNORECASE,
)
embedding = await generate_embedding(text, db=db, tenant_id=tenant_id)
if not embedding:
return False