LLM Post-Training, Agentic RL, Self-Evolving Agents and Sales Agents

Overview:

We research on efficient LLM post-training methods for agentic-RL and self-evolving agents, including error-driven learning from failed trajectories, fine-grained credit assignment, progressive curricula for long-horizon RL, and router-aware MoE fine-tuning; also built SOP-guided RL for conversational sales agents and realistic user simulation for interactive agents. Research has been translating into production systems for financial marketing, customer service, and compliance.

Publications: