feat(B-LLM): Zentraler LLM Client — llm_complete() + llm_embed() + Migration + Tests + Doku
Check Cross-Plugin Imports / check (push) Has been cancelled
Check Cross-Plugin Imports / check (push) Has been cancelled
B-LLM: llm_client.py um generische llm_complete() und llm_embed() erweitert - Provider-Auswahl, API-Key-Auflösung, Error-Handling, Cost-Tracking - Retry mit Exponential-Backoff für transient errors - Timeout konfigurierbar - Helper: get_api_credentials(), build_model(), _classify_error() B-LLM-MIG: Alle 8 direkten litellm.acompletion() Calls auf llm_complete() umgestellt - agent_runner.py, query_understanding.py (2x), ai_proactive (3x), ai_assistant (2x) - 0 verbleibende direkte litellm.acompletion() Calls außerhalb llm_client.py B-LLM-TEST: 39 Tests in test_llm_client.py — alle grün - Mock mode, error handling, embed, helpers, backward compat B-LLM-DOC: Plugin-Dev-Guide Kapitel 7 (LLM Integration) hinzugefügt
This commit is contained in:
@@ -1,7 +1,7 @@
|
||||
"""AI Participant Handler — bridges the kommunikation plugin with the AI Assistant.
|
||||
|
||||
When a message is received in a conversation that includes the 'ai' participant,
|
||||
this handler generates an LLM response using litellm.acompletion (non-streaming)
|
||||
this handler generates an LLM response using llm_complete (non-streaming)
|
||||
and returns it as a new message in the conversation.
|
||||
"""
|
||||
|
||||
@@ -11,8 +11,7 @@ import logging
|
||||
import uuid
|
||||
from typing import Any
|
||||
|
||||
import litellm
|
||||
|
||||
from app.ai.llm_client import llm_complete
|
||||
from app.core.db import create_db_session
|
||||
from app.plugins.builtins.kommunikation.contracts import ParticipantHandler
|
||||
|
||||
@@ -136,11 +135,18 @@ class AIParticipantHandler(ParticipantHandler):
|
||||
if provider.base_url:
|
||||
params["api_base"] = provider.base_url
|
||||
|
||||
# Ensure non-streaming for acompletion
|
||||
params["stream"] = False
|
||||
# Ensure non-streaming
|
||||
params.pop("stream", None)
|
||||
|
||||
response = await litellm.acompletion(**params)
|
||||
response_text = response.choices[0].message.content or ""
|
||||
result = await llm_complete(
|
||||
model=params.get("model", "gpt-4o-mini"),
|
||||
messages=params.get("messages", []),
|
||||
temperature=params.get("temperature", 0.7),
|
||||
max_tokens=params.get("max_tokens", 2048),
|
||||
api_key=params.get("api_key"),
|
||||
api_base=params.get("api_base"),
|
||||
)
|
||||
response_text = result["content"] or ""
|
||||
|
||||
if not response_text.strip():
|
||||
response_text = "*(keine Antwort generiert)*"
|
||||
|
||||
Reference in New Issue
Block a user