Glossary

Glossary

Quick reference for the core terminology used across the LinkSoul AgentSDK docs, grouped into business concepts; passive vs. active interaction in detail; SDK terms; model / audio-video terms; dialogue policy / reject; and JSON protocol fields. Each entry is one sentence + a pointer to the relevant doc page.

1. Business concepts

TermDescription
LinkSoul Open PlatformThe agibot developer-facing platform for app registration, agent configuration, and gateway access. appId / appKey / appSecret are issued here.
LinkskyGatewayThe open-platform ingress gateway at wss://open.agibot.com/api/V1/open-portal/app/wss/agent-sdk relaying bidirectional WebSocket messages between SDK and robots.
AgentA virtual character configured on LinkSoul (persona, skills, knowledge base). Robots bind to an agent on-device; once online the SDK takes over its semantics / skill / task-flow handling.
Custom development (二开)Building your own ASR / LLM / skill / task-flow logic on top of the default LinkSoul capabilities via this SDK.
RobotThe physical device. Bound to one or more agents, it acts as both message source (audio / video / text) and executor (TTS playback, skill actions).
Passive interactionRobot initiates → SDK processes → SDK fills the Response (ASR / LLM / VLM / TTS / skill / interrupt).
Active interactionSDK initiates toward the robot. v1.4.0 only exposes Task Flow publicly.
Task FlowSDK-driven multi-step task orchestration with start / task / end / interrupt phases; payload is a JSON array; SIMPLE and COMPLEX modes supported.
SkillA named action the robot can execute (e.g. move_forward, wave_hands), dispatched via response.onSkill(...) or task flow. Full catalogue: Semantic Skills.
Greet"Active greeting" scenario triggered by face / UID / video / business signal; SDK fills VLM/TTS via onGreetSignal + GreetResponse.
Interrupt / Barge-inWhen the user speaks again during robot playback, the SDK calls response.onInterrupt(...) to terminate the current playback and switch to the new intent.

2. Passive vs. active interaction in detail

The fundamental difference is who initiates the message:

DimensionPassiveActive
InitiatorRobot (device side) pushes via LinkskyGatewaySDK (your code) dispatches to the robot
TriggerUser speaks / face detected / video frame / device-side state pushBusiness system decides the robot should run a task (e.g. "turn on living-room light", "patrol 5 waypoints")
Entry API8 register* methods (registerAudio2Llm / registerAsr2Llm / registerAsrVideo2Vlm …) + matching onRequest callbacksagentSdk.registerTaskFlow(TaskFlowRequest) + flow.startRequest / taskRequest / endRequest / interruptRequest
Data carrierCallback args text / buf / video / param + the Response objectTaskFlowRequest payload (JSON array) + two-phase RequestAck / ExecuteAck
Reply pathInside onRequest call response.onLlm* / onTts* / onSkill / onInterrupt / onErrorEach phase attaches FlowStart/Task/End/InterruptCallback; gateway / robot replies asynchronously
v1.4.0 surfaceAll 8 passive callbacks availableOnly task flow is public; the rest of the active surface (skill/state query, pull/push stream) was removed from public API

Passive

  • Initiator: the robot. The user speaks at it, a face appears, device state changes; the device packages a protocol event and the gateway pushes it to the SDK.
  • SDK routing: the SDK dispatches the event to the right *Callback.onRequest(...). The 8 passive callbacks are named after their input → output pair:
    InputOutputCallback
    AudioLLM textAudio2LlmCallback
    AudioLLM + TTSAudio2TtsCallback
    ASR textLLM textAsr2LlmCallback
    ASR textLLM + TTSAsr2TtsCallback
    ASR text + videoVLM textAsrVideo2VlmCallback
    ASR text + videoVLM + TTSAsrVideo2TtsCallback
    Audio + videoVLM textAudioVideo2VlmCallback
    Audio + videoVLM + TTSAudioVideo2TtsCallback
  • One passive callback type per SDK instance — enforced by AgentSdk.
  • Shared PassiveCallback base carries device-level events that are independent of the input/output combination: onRobotOnline / onRobotOffline / onFaceInfo / onVideoFrame / onGreetSignal / onState. They run alongside onRequest, not inside it.
  • Replies must be filled in explicitly — for reject, just return; for a normal reply, follow the order onInterrupt → onSkill?(optional) → onLlm*/onTts*/onVlm* → onLlmDone/onTtsDone.
  • See Passive Callbacks Guide.

Active — task flow only in v1.4.0

  • Initiator: your business code. When the upper layer (external trigger, scheduler, another model) decides the robot should actively run a task, the SDK dispatches it to the robot through LinkskyGateway.
  • The only public entry: registerTaskFlow(TaskFlowRequest). Skill query / state query / pull stream / push-listen — exposed in v1.3.0 but never functional — were removed from the public API in v1.4.0.
  • Four phases:
    1. startRequest(payload, callback, timeout) — start a task flow.
    2. taskRequest(payload, callback, timeout) — send a step (callable multiple times).
    3. endRequest(callback, timeout) — normal end.
    4. interruptRequest(callback, timeout) — abort the current flow.
  • Two-phase ack: every *Request produces RequestAck (robot received) and ExecuteAck (robot finished); code=0 per phase = success.
  • Payload shape: JSON array (COMPLEX) or single-step object (SIMPLE), described by action_group / action_type etc. See Task-flow payload protocol.
  • State feedback: while the flow runs, the robot's device-side state changes (navigation progress, playback status …) come back through the passive channel's onState(...). Active and passive channels meet at the state layer.
  • See Active Operations Guide.

Passive ↔ active collaboration

Real projects combine both channels:

  • Passive triggers active: user says "take me to the meeting room" → Asr2LlmCallback.onRequest fires → business code parses navigation intent → dispatches a TaskFlowRequest.
  • Passive interrupts active: while a flow is running the user speaks again → passive channel receives a new eventId onRequest → business code calls flow.interruptRequest(...).
  • State bridges them: a task_status update arriving via passive onState drives the active side's next decision.

3. SDK terms

TermDescription
appIdUnique application identifier issued by the LinkSoul open platform.
appKeyNew in v1.4.0 — access key used in HMAC-SHA256 signing alongside appSecret.
appSecretHMAC-SHA256 signing key. Never ship it to clients, logs, or git.
agentIdThe agent ID configured on LinkSoul (e.g. AGENT_0000001). Robots bind to an agent on-device; the SDK delivers it via onRobotOnline(agentId, ...). Not a hardware identifier.
eventIdUnique id for one full conversation / event; every response must carry it back.
itemIdIdentifier for multiple intents / response chunks within one eventId (multiple LLM segments, multiple TTS pieces, etc.).
flowIdUnique id for a task-flow instance, generated by IdGenerator.generateFlowId().
signalIdUnique id for a business signal (e.g. greet).
flagLifecycle flag exclusive to audio/video callbacks: START / APPEND / COMMIT etc.
AgentSdkSDK main entry class (singleton — same appId always returns the same instance). Owns connection, auth, callback registration, and task-flow dispatch.
AgentParamType-safe key-value parameter container with chained setters (setString / setInteger / setDouble / setBoolean / setLong / setShort / setObject / setList).
AgentMetaRobot metadata returned by onRobotOnline: wakeup word, city, robot name, buzz word, and custom fields.
PassiveCallbackBase class of every passive callback (Audio2Llm / Asr2Llm / AudioVideo2Vlm …). Provides shared onRobotOnline / onRobotOffline / onFaceInfo / onVideoFrame / onGreetSignal / onState.
Response objectThe reply object handed to each onRequest. Provides onAsr* / onLlm* / onVlm* / onTts* / onSkill / onInterrupt / onError depending on mode.
onState(stateName, stateValue)Robot-side state push. The 22 stateName modules are documented in Robot-side state modules.
RequestAck / ExecuteAckTwo-phase task-flow acknowledgement — request received vs. execution finished; code=0 means success.

4. Model / audio-video terms

TermDescription
ASR (Automatic Speech Recognition)Audio → text. Performed by the gateway or upstream component; usually arrives as a text field.
NLU (Natural Language Understanding)Intent classification / slot filling / entity extraction over the ASR text. Custom-development integrators usually plug in their own NLU on top of onRequest's text argument and then decide between LLM chat, skill dispatch, or reject.
LLM (Large Language Model)Text understanding and generation. SDK streams it back via onLlmItemDelta/Done.
VLM (Vision-Language Model)Multimodal model accepting image/video + text; surfaced through the onVlm* family.
TTS (Text-To-Speech)Text → audio, streamed back as base64 audio chunks via onTtsDelta/Done.
VAD (Voice Activity Detection)Detects user speech start/stop, mapping to protocol events agentsdk.audio_request.start / commit.
RAG knowledge base (Retrieval-Augmented Generation)Retrieve relevant chunks from a vector / lexical knowledge base, feed them as context to the LLM, then generate the answer. Used to ground custom LLM flows on private docs / FAQs and reduce hallucinations. In the SDK it usually sits between receiving text in onRequest and the first onLlmItemDelta.
Key frame / isKeyFrameMarks whether a video-frame callback is an H.264 IDR key frame.
base64 audio chunkTTS / greet-TTS streaming chunks are wrapped as base64 strings in the protocol payload.

5. Dialogue policy & reject

TermDescription
RejectThe dialogue system, after ASR / NLU / LLM, decides the input "should not be handled" and either stays silent or returns a fallback. In the SDK this maps to onRequest returning without calling any response.on*, or returning an error only.
General rejectDomain-agnostic reject. Typical triggers: empty text / very-low ASR confidence / noise / meaningless filler / global sensitive-word hits. Usually the first filter, run before any vertical logic.
Domain rejectReject specific to a business vertical (medical, finance, household, child companion, etc.) for inputs out of scope or explicitly forbidden in that domain. Driven by the vertical's intent classifier confidence, allow/deny lists, or domain-specific sensitive-word tables; on hit, stay silent or return a vertical-specific fallback.
FallbackThe fixed reply used when reject hits or LLM fails (e.g. "Sorry, I can't answer that right now").

6. JSON protocol fields (task flow / state)

FieldDescription
action_groupAction group inside a task-flow payload; defines a parallel/serial bundle of sub-actions.
action_typeAction type (e.g. target_poi, tts_speak); defined by skills on the LinkSoul platform. See Task-flow payload protocol.
stateName (22 modules)Robot-side state names: robot_pose, tts_status, robot_form, task_status, etc. Field schemas in Robot-side state modules.
SIMPLE / COMPLEX modesTwo task-flow orchestration modes: SIMPLE — single-shot send; COMPLEX — full three-phase (start → task* → end), interruptible mid-flight.

Related entry points