基于离散标签化的机器人动作策略优化

Optimization of Robot Action Strategies Based on Discrete Labeling

  • 摘要: 针对视觉语言动作(VLA)模型受推理效率制约,仅能生成单步动作、无法利用历史动作经验,且单步偏差易导致机器人作业任务失败的问题,本文旨在提升模型推理效率,实现多步轨迹高效生成,并借助多步轨迹优化当前动作。本文基于OpenVLA-7B模型,构建了一种基于残差VQ-VAE离散化标签方式,通过分别训练训练编解码器和构建向量代码本的方式,有效压缩机器人动作的高维连续信息。采用离散化标签的方式能够有效降低模型训练的复杂度,提高推理效率。基于生成的动作序列,提出了一种应用前序动作序列的动作优化方法,以充分利用前序经验,提高模型在执行过程中的稳定性。在LIBERO基准和实际场景中对改进的模型进行了验证,改进后的模型成功率和推理效率均优于基座模型,且推理效率提升78.6%,证明了优化方法的有效性。

     

    Abstract: Current Vision-Language-Action (VLA) models suffer from low inference efficiency, resulting in exclusive single-step action generation, inability to use prior action experiences, and frequent failures of robotic manipulation tasks caused by single-step deviations. We target the enhancement of inference efficiency, efficient multi-step trajectory generation and trajectory-based action optimization. Built upon the OpenVLA-7B model, a Residual VQ-VAE based discrete labeling scheme is developed. Separate training of the encoder-decoder and construction of a vector codebook effectively compress the high-dimensional continuous information of robotic actions. Discrete labels greatly lower training complexity and raise inference efficiency. Leveraging generated action sequences, an optimization method for ctions is presented to exploit preceding action experiences and strengthen model stability during execution. Evaluated on the LIBERO benchmark and in real-world environments, the improved model outperforms the baseline model in both success rate and inference efficiency, with a 78.6% increase in inference efficiency, which confirms the validity of the proposed method.

     

/

返回文章
返回