Glossary

Glossary

Quick reference for the core terminology used across the LinkSoul AgentSDK (Python) docs, grouped into business concepts; passive vs. active interaction in detail; SDK terms; model / audio-video terms; dialogue policy / reject; and JSON protocol fields. Each entry is one sentence + a pointer to the relevant doc page.

1. Business concepts

TermDescription
LinkSoul Open PlatformThe agibot developer-facing platform for app registration, agent configuration, and gateway access. app_id / app_key / app_secret are issued here.
LinkskyGatewayThe open-platform ingress gateway at wss://open.agibot.com/api/V1/open-portal/app/wss/agent-sdk relaying bidirectional WebSocket messages between SDK and robots.
AgentA virtual character configured on LinkSoul (persona, skills, knowledge base). Robots bind to an agent on-device; once online the SDK takes over its semantics / skill / task-flow handling.
Custom development (二开)Building your own ASR / LLM / skill / task-flow logic on top of the default LinkSoul capabilities via this SDK.
RobotThe physical device. Bound to one or more agents, it acts as both message source (audio / video / text) and executor (TTS playback, skill actions).
Passive interactionRobot initiates → SDK processes → SDK fills the Response (ASR / LLM / VLM / TTS / skill / interrupt).
Active interactionSDK initiates toward the robot. v1.4.0 only exposes Task Flow publicly.
Task FlowSDK-driven multi-step task orchestration with start / task / end / interrupt phases; payload is a JSON array; SIMPLE and COMPLEX modes supported.
SkillA named action the robot can execute (e.g. move_forward, wave_hands), dispatched via response.on_skill(...) or task flow. Full catalogue: Semantic Skills.
Greet"Active greeting" scenario triggered by face / UID / video / business signal; SDK fills VLM/TTS via on_greet_signal + GreetResponse.
Interrupt / Barge-inWhen the user speaks again during robot playback, the SDK calls response.on_interrupt(...) to terminate the current playback and switch to the new intent.

2. Passive vs. active interaction in detail

The fundamental difference is who initiates the message:

DimensionPassiveActive
InitiatorRobot (device side) pushes via LinkskyGatewaySDK (your code) dispatches to the robot
TriggerUser speaks / face detected / video frame / device-side state pushBusiness system decides the robot should run a task (e.g. "turn on living-room light", "patrol 5 waypoints")
Entry API8 register_* methods (register_audio2_llm / register_asr2_llm / register_asr_video2_vlm …) + matching on_request callbacksagent_sdk.register_task_flow(TaskFlowRequest) + flow.start_request / task_request / end_request / interrupt_request
Data carrierCallback args text / buf / video / param + the Response objectTaskFlowRequest payload (JSON array) + two-phase RequestAck / ExecuteAck
Reply pathInside on_request call response.on_llm_* / on_tts_* / on_skill / on_interrupt / on_errorEach phase attaches FlowStart/Task/End/InterruptCallback; gateway / robot replies asynchronously
v1.4.0 surfaceAll 8 passive callbacks availableOnly task flow is public; the rest of the active surface (skill/state query, pull/push stream) was removed from public API

Passive

  • Initiator: the robot. The user speaks at it, a face appears, device state changes; the device packages a protocol event and the gateway pushes it to the SDK.
  • SDK routing: the SDK dispatches the event to the right *Callback.on_request(...). The 8 passive callbacks are named after their input → output pair:
    InputOutputCallback
    AudioLLM textAudio2LlmCallback
    AudioLLM + TTSAudio2TtsCallback
    ASR textLLM textAsr2LlmCallback
    ASR textLLM + TTSAsr2TtsCallback
    ASR text + videoVLM textAsrVideo2VlmCallback
    ASR text + videoVLM + TTSAsrVideo2TtsCallback
    Audio + videoVLM textAudioVideo2VlmCallback
    Audio + videoVLM + TTSAudioVideo2TtsCallback
  • One passive callback type per SDK instance — enforced by AgentSdk.
  • Shared PassiveCallback base carries device-level events that are independent of the input/output combination: on_robot_online / on_robot_offline / on_face_info / on_video_frame / on_greet_signal / on_state. They run alongside on_request, not inside it.
  • Replies must be filled in explicitly — for reject, just return; for a normal reply, follow the order on_interrupt → on_skill?(optional) → on_llm_*/on_tts_*/on_vlm_* → on_llm_done/on_tts_done.
  • See Passive Callbacks Guide.

Active — task flow only in v1.4.0

  • Initiator: your business code. When the upper layer (external trigger, scheduler, another model) decides the robot should actively run a task, the SDK dispatches it to the robot through LinkskyGateway.
  • The only public entry: register_task_flow(TaskFlowRequest). Skill query / state query / pull stream / push-listen — exposed in v1.3.0 but never functional — were removed from the public API in v1.4.0.
  • Four phases:
    1. start_request(payload, callback, timeout) — start a task flow.
    2. task_request(payload, callback, timeout) — send a step (callable multiple times).
    3. end_request(callback, timeout) — normal end.
    4. interrupt_request(callback, timeout) — abort the current flow.
  • Two-phase ack: every *_request produces RequestAck (robot received) and ExecuteAck (robot finished); code=0 per phase = success.
  • Payload shape: JSON array (COMPLEX) or single-step object (SIMPLE), described by action_group / action_type etc. See Task-flow payload protocol.
  • State feedback: while the flow runs, the robot's device-side state changes (navigation progress, playback status …) come back through the passive channel's on_state(...). Active and passive channels meet at the state layer.
  • See Active Operations Guide.

Passive ↔ active collaboration

Real projects combine both channels:

  • Passive triggers active: user says "take me to the meeting room" → Asr2LlmCallback.on_request fires → business code parses navigation intent → dispatches a TaskFlowRequest.
  • Passive interrupts active: while a flow is running the user speaks again → passive channel receives a new event_id on_request → business code calls flow.interrupt_request(...).
  • State bridges them: a task_status update arriving via passive on_state drives the active side's next decision.

3. SDK terms

TermDescription
app_idUnique application identifier issued by the LinkSoul open platform.
app_keyNew in v1.4.0 — access key used in HMAC-SHA256 signing alongside app_secret.
app_secretHMAC-SHA256 signing key. Never ship it to clients, logs, or git.
agent_idThe agent ID configured on LinkSoul (e.g. AGENT_0000001). Robots bind to an agent on-device; the SDK delivers it via on_robot_online(agent_id, ...). Not a hardware identifier.
event_idUnique id for one full conversation / event; every response must carry it back.
item_idIdentifier for multiple intents / response chunks within one event_id (multiple LLM segments, multiple TTS pieces, etc.).
flow_idUnique id for a task-flow instance, generated by IdGenerator.generate_flow_id().
signal_idUnique id for a business signal (e.g. greet).
flagLifecycle flag exclusive to audio/video callbacks: START / APPEND / COMMIT etc.
AgentSdkSDK main entry class (singleton — same app_id always returns the same instance). Owns connection, auth, callback registration, and task-flow dispatch.
AgentParamType-safe key-value parameter container with chained setters (set_string / set_integer / set_double / set_boolean / set_long / set_object / set_list).
AgentMetaRobot metadata returned by on_robot_online: wakeup word, city, robot name, buzz word, and custom fields.
PassiveCallbackBase class of every passive callback (Audio2LlmCallback / Asr2LlmCallback / AudioVideo2VlmCallback …). Provides shared on_robot_online / on_robot_offline / on_face_info / on_video_frame / on_greet_signal / on_state.
Response objectThe reply object handed to each on_request. Provides on_asr_* / on_llm_* / on_vlm_* / on_tts_* / on_skill / on_interrupt / on_error depending on mode.
on_state(state_name, state_value)Robot-side state push. The 22 state_name modules are documented in Robot-side state modules.
RequestAck / ExecuteAckTwo-phase task-flow acknowledgement — request received vs. execution finished; code=0 means success.

4. Model / audio-video terms

TermDescription
ASR (Automatic Speech Recognition)Audio → text. Performed by the gateway or upstream component; usually arrives as a text field.
NLU (Natural Language Understanding)Intent classification / slot filling / entity extraction over the ASR text. Custom-development integrators usually plug in their own NLU on top of on_request's text argument and then decide between LLM chat, skill dispatch, or reject.
LLM (Large Language Model)Text understanding and generation. SDK streams it back via on_llm_item_delta/done.
VLM (Vision-Language Model)Multimodal model accepting image/video + text; surfaced through the on_vlm_* family.
TTS (Text-To-Speech)Text → audio, streamed back as base64 audio chunks via on_tts_delta/done.
VAD (Voice Activity Detection)Detects user speech start/stop, mapping to protocol events agentsdk.audio_request.start / commit.
RAG knowledge base (Retrieval-Augmented Generation)Retrieve relevant chunks from a vector / lexical knowledge base, feed them as context to the LLM, then generate the answer. Used to ground custom LLM flows on private docs / FAQs and reduce hallucinations. In the SDK it usually sits between receiving text in on_request and the first on_llm_item_delta.
Key frame / is_key_frameMarks whether a video-frame callback is an H.264 IDR key frame.
base64 audio chunkTTS / greet-TTS streaming chunks are wrapped as base64 strings in the protocol payload.

5. Dialogue policy & reject

TermDescription
RejectThe dialogue system, after ASR / NLU / LLM, decides the input "should not be handled" and either stays silent or returns a fallback. In the SDK this maps to on_request returning without calling any response.on_*, or returning an error only.
General rejectDomain-agnostic reject. Typical triggers: empty text / very-low ASR confidence / noise / meaningless filler / global sensitive-word hits. Usually the first filter, run before any vertical logic.
Domain rejectReject specific to a business vertical (medical, finance, household, child companion, etc.) for inputs out of scope or explicitly forbidden in that domain. Driven by the vertical's intent classifier confidence, allow/deny lists, or domain-specific sensitive-word tables; on hit, stay silent or return a vertical-specific fallback.
FallbackThe fixed reply used when reject hits or LLM fails (e.g. "Sorry, I can't answer that right now").

6. JSON protocol fields (task flow / state)

FieldDescription
action_groupAction group inside a task-flow payload; defines a parallel/serial bundle of sub-actions.
action_typeAction type (e.g. target_poi, tts_speak); defined by skills on the LinkSoul platform. See Task-flow payload protocol.
state_name (22 modules)Robot-side state names: robot_pose, tts_status, robot_form, task_status, etc. Field schemas in Robot-side state modules.
SIMPLE / COMPLEX modesTwo task-flow orchestration modes: SIMPLE — single-shot send; COMPLEX — full three-phase (start → task* → end), interruptible mid-flight.

Related entry points