Glossary
Glossary
Quick reference for the core terminology used across the LinkSoul AgentSDK (Python) docs, grouped into business concepts; passive vs. active interaction in detail; SDK terms; model / audio-video terms; dialogue policy / reject; and JSON protocol fields. Each entry is one sentence + a pointer to the relevant doc page.
1. Business concepts
| Term | Description |
|---|---|
| LinkSoul Open Platform | The agibot developer-facing platform for app registration, agent configuration, and gateway access. app_id / app_key / app_secret are issued here. |
| LinkskyGateway | The open-platform ingress gateway at wss://open.agibot.com/api/V1/open-portal/app/wss/agent-sdk relaying bidirectional WebSocket messages between SDK and robots. |
| Agent | A virtual character configured on LinkSoul (persona, skills, knowledge base). Robots bind to an agent on-device; once online the SDK takes over its semantics / skill / task-flow handling. |
| Custom development (二开) | Building your own ASR / LLM / skill / task-flow logic on top of the default LinkSoul capabilities via this SDK. |
| Robot | The physical device. Bound to one or more agents, it acts as both message source (audio / video / text) and executor (TTS playback, skill actions). |
| Passive interaction | Robot initiates → SDK processes → SDK fills the Response (ASR / LLM / VLM / TTS / skill / interrupt). |
| Active interaction | SDK initiates toward the robot. v1.4.0 only exposes Task Flow publicly. |
| Task Flow | SDK-driven multi-step task orchestration with start / task / end / interrupt phases; payload is a JSON array; SIMPLE and COMPLEX modes supported. |
| Skill | A named action the robot can execute (e.g. move_forward, wave_hands), dispatched via response.on_skill(...) or task flow. Full catalogue: Semantic Skills. |
| Greet | "Active greeting" scenario triggered by face / UID / video / business signal; SDK fills VLM/TTS via on_greet_signal + GreetResponse. |
| Interrupt / Barge-in | When the user speaks again during robot playback, the SDK calls response.on_interrupt(...) to terminate the current playback and switch to the new intent. |
2. Passive vs. active interaction in detail
The fundamental difference is who initiates the message:
| Dimension | Passive | Active |
|---|---|---|
| Initiator | Robot (device side) pushes via LinkskyGateway | SDK (your code) dispatches to the robot |
| Trigger | User speaks / face detected / video frame / device-side state push | Business system decides the robot should run a task (e.g. "turn on living-room light", "patrol 5 waypoints") |
| Entry API | 8 register_* methods (register_audio2_llm / register_asr2_llm / register_asr_video2_vlm …) + matching on_request callbacks | agent_sdk.register_task_flow(TaskFlowRequest) + flow.start_request / task_request / end_request / interrupt_request |
| Data carrier | Callback args text / buf / video / param + the Response object | TaskFlowRequest payload (JSON array) + two-phase RequestAck / ExecuteAck |
| Reply path | Inside on_request call response.on_llm_* / on_tts_* / on_skill / on_interrupt / on_error | Each phase attaches FlowStart/Task/End/InterruptCallback; gateway / robot replies asynchronously |
| v1.4.0 surface | All 8 passive callbacks available | Only task flow is public; the rest of the active surface (skill/state query, pull/push stream) was removed from public API |
Passive
- Initiator: the robot. The user speaks at it, a face appears, device state changes; the device packages a protocol event and the gateway pushes it to the SDK.
- SDK routing: the SDK dispatches the event to the right
*Callback.on_request(...). The 8 passive callbacks are named after their input → output pair:Input Output Callback Audio LLM text Audio2LlmCallbackAudio LLM + TTS Audio2TtsCallbackASR text LLM text Asr2LlmCallbackASR text LLM + TTS Asr2TtsCallbackASR text + video VLM text AsrVideo2VlmCallbackASR text + video VLM + TTS AsrVideo2TtsCallbackAudio + video VLM text AudioVideo2VlmCallbackAudio + video VLM + TTS AudioVideo2TtsCallback - One passive callback type per SDK instance — enforced by
AgentSdk. - Shared
PassiveCallbackbase carries device-level events that are independent of the input/output combination:on_robot_online/on_robot_offline/on_face_info/on_video_frame/on_greet_signal/on_state. They run alongsideon_request, not inside it. - Replies must be filled in explicitly — for reject, just
return; for a normal reply, follow the orderon_interrupt → on_skill?(optional) → on_llm_*/on_tts_*/on_vlm_* → on_llm_done/on_tts_done. - See Passive Callbacks Guide.
Active — task flow only in v1.4.0
- Initiator: your business code. When the upper layer (external trigger, scheduler, another model) decides the robot should actively run a task, the SDK dispatches it to the robot through
LinkskyGateway. - The only public entry:
register_task_flow(TaskFlowRequest). Skill query / state query / pull stream / push-listen — exposed in v1.3.0 but never functional — were removed from the public API in v1.4.0. - Four phases:
start_request(payload, callback, timeout)— start a task flow.task_request(payload, callback, timeout)— send a step (callable multiple times).end_request(callback, timeout)— normal end.interrupt_request(callback, timeout)— abort the current flow.
- Two-phase ack: every
*_requestproducesRequestAck(robot received) andExecuteAck(robot finished);code=0per phase = success. - Payload shape: JSON array (COMPLEX) or single-step object (SIMPLE), described by
action_group/action_typeetc. See Task-flow payload protocol. - State feedback: while the flow runs, the robot's device-side state changes (navigation progress, playback status …) come back through the passive channel's
on_state(...). Active and passive channels meet at the state layer. - See Active Operations Guide.
Passive ↔ active collaboration
Real projects combine both channels:
- Passive triggers active: user says "take me to the meeting room" →
Asr2LlmCallback.on_requestfires → business code parses navigation intent → dispatches aTaskFlowRequest. - Passive interrupts active: while a flow is running the user speaks again → passive channel receives a new
event_idon_request→ business code callsflow.interrupt_request(...). - State bridges them: a
task_statusupdate arriving via passiveon_statedrives the active side's next decision.
3. SDK terms
| Term | Description |
|---|---|
app_id | Unique application identifier issued by the LinkSoul open platform. |
app_key | New in v1.4.0 — access key used in HMAC-SHA256 signing alongside app_secret. |
app_secret | HMAC-SHA256 signing key. Never ship it to clients, logs, or git. |
agent_id | The agent ID configured on LinkSoul (e.g. AGENT_0000001). Robots bind to an agent on-device; the SDK delivers it via on_robot_online(agent_id, ...). Not a hardware identifier. |
event_id | Unique id for one full conversation / event; every response must carry it back. |
item_id | Identifier for multiple intents / response chunks within one event_id (multiple LLM segments, multiple TTS pieces, etc.). |
flow_id | Unique id for a task-flow instance, generated by IdGenerator.generate_flow_id(). |
signal_id | Unique id for a business signal (e.g. greet). |
flag | Lifecycle flag exclusive to audio/video callbacks: START / APPEND / COMMIT etc. |
AgentSdk | SDK main entry class (singleton — same app_id always returns the same instance). Owns connection, auth, callback registration, and task-flow dispatch. |
AgentParam | Type-safe key-value parameter container with chained setters (set_string / set_integer / set_double / set_boolean / set_long / set_object / set_list). |
AgentMeta | Robot metadata returned by on_robot_online: wakeup word, city, robot name, buzz word, and custom fields. |
PassiveCallback | Base class of every passive callback (Audio2LlmCallback / Asr2LlmCallback / AudioVideo2VlmCallback …). Provides shared on_robot_online / on_robot_offline / on_face_info / on_video_frame / on_greet_signal / on_state. |
| Response object | The reply object handed to each on_request. Provides on_asr_* / on_llm_* / on_vlm_* / on_tts_* / on_skill / on_interrupt / on_error depending on mode. |
on_state(state_name, state_value) | Robot-side state push. The 22 state_name modules are documented in Robot-side state modules. |
RequestAck / ExecuteAck | Two-phase task-flow acknowledgement — request received vs. execution finished; code=0 means success. |
4. Model / audio-video terms
| Term | Description |
|---|---|
| ASR (Automatic Speech Recognition) | Audio → text. Performed by the gateway or upstream component; usually arrives as a text field. |
| NLU (Natural Language Understanding) | Intent classification / slot filling / entity extraction over the ASR text. Custom-development integrators usually plug in their own NLU on top of on_request's text argument and then decide between LLM chat, skill dispatch, or reject. |
| LLM (Large Language Model) | Text understanding and generation. SDK streams it back via on_llm_item_delta/done. |
| VLM (Vision-Language Model) | Multimodal model accepting image/video + text; surfaced through the on_vlm_* family. |
| TTS (Text-To-Speech) | Text → audio, streamed back as base64 audio chunks via on_tts_delta/done. |
| VAD (Voice Activity Detection) | Detects user speech start/stop, mapping to protocol events agentsdk.audio_request.start / commit. |
| RAG knowledge base (Retrieval-Augmented Generation) | Retrieve relevant chunks from a vector / lexical knowledge base, feed them as context to the LLM, then generate the answer. Used to ground custom LLM flows on private docs / FAQs and reduce hallucinations. In the SDK it usually sits between receiving text in on_request and the first on_llm_item_delta. |
Key frame / is_key_frame | Marks whether a video-frame callback is an H.264 IDR key frame. |
| base64 audio chunk | TTS / greet-TTS streaming chunks are wrapped as base64 strings in the protocol payload. |
5. Dialogue policy & reject
| Term | Description |
|---|---|
| Reject | The dialogue system, after ASR / NLU / LLM, decides the input "should not be handled" and either stays silent or returns a fallback. In the SDK this maps to on_request returning without calling any response.on_*, or returning an error only. |
| General reject | Domain-agnostic reject. Typical triggers: empty text / very-low ASR confidence / noise / meaningless filler / global sensitive-word hits. Usually the first filter, run before any vertical logic. |
| Domain reject | Reject specific to a business vertical (medical, finance, household, child companion, etc.) for inputs out of scope or explicitly forbidden in that domain. Driven by the vertical's intent classifier confidence, allow/deny lists, or domain-specific sensitive-word tables; on hit, stay silent or return a vertical-specific fallback. |
| Fallback | The fixed reply used when reject hits or LLM fails (e.g. "Sorry, I can't answer that right now"). |
6. JSON protocol fields (task flow / state)
| Field | Description |
|---|---|
action_group | Action group inside a task-flow payload; defines a parallel/serial bundle of sub-actions. |
action_type | Action type (e.g. target_poi, tts_speak); defined by skills on the LinkSoul platform. See Task-flow payload protocol. |
state_name (22 modules) | Robot-side state names: robot_pose, tts_status, robot_form, task_status, etc. Field schemas in Robot-side state modules. |
| SIMPLE / COMPLEX modes | Two task-flow orchestration modes: SIMPLE — single-shot send; COMPLEX — full three-phase (start → task* → end), interruptible mid-flight. |
Related entry points
- Quick Start — 5-minute onboarding
- Architecture — threading model, reconnect mechanism
- Passive Callbacks Guide — 8 passive callback types in detail
- Active Operations Guide — task-flow lifecycle
- Task-flow payload protocol —
action_group/action_type - Semantic Skills — full skill catalogue
- Robot-side state modules — 22
state_namemodules