LG Partners with Nvidia on 10,000-Square-Meter Robot Data Factory Targeting 100,000 Training Hours

LG Electronics hosted senior Nvidia leadership at its Yangjae R&D Campus in Seoul on August 18, 2026, advancing a joint initiative to build physical AI training pipelines and target 100,000 hours of embodied robotics data by the end of the year. The site review took place five days after LG Group and Nvidia signed a strategic memorandum of understanding at Nvidia headquarters in Santa Clara on August 13. The accelerated timeline reflects LG's effort to convert its industrial manufacturing infra

2 min
LG Partners with Nvidia on 10,000-Square-Meter Robot Data Factory Targeting 100,000 Training Hours

LG Electronics hosted senior Nvidia leadership at its Yangjae R&D Campus in Seoul on August 18, 2026, advancing a joint initiative to build physical AI training pipelines and target 100,000 hours of embodied robotics data by the end of the year.

The site review took place five days after LG Group and Nvidia signed a strategic memorandum of understanding at Nvidia headquarters in Santa Clara on August 13. The accelerated timeline reflects LG's effort to convert its industrial manufacturing infrastructure and consumer appliance operations into a continuous data engine for general-purpose robotic systems.

Inside the 10,000-Square-Meter Facility

The Yangjae Data Factory spans four floors (basement level 1 through the third floor) covering 10,000 square meters. LG plans to deploy hundreds of robots across modular testing environments within the building by the end of 2026.

The facility functions as a physical-to-digital training ground divided into specialized task domains:

  • Domestic Environments: LG CLOiD home service robots execute continuous household maintenance, cleaning, and manipulation tasks on automated loops.
  • Industrial Manufacturing: Assembly cells modeled after LG's washing machine manufacturing facility in Clarksville, Tennessee test part handling, sorting, and mechanical assembly.
  • Logistics and Manipulation: LG CNS operates automated logistics workflows, while LG Innotek trains multi-fingered robotic hands on precision dexterity tasks.
LG and Nvidia Embodied AI Data Flywheel Pipeline

The 100,000-Hour Training Objective

Physical data captured from sensors, joint encoders, and vision systems across the Yangjae facility is ingested into Nvidia's robotics computing stack. LG uses Nvidia Cosmos world foundation models, Omniverse digital twins, and the Isaac robotics platform to filter, augment, and synthesize the captured trajectories into simulation.

Combining real-world physical capture with generative synthetic data, LG aims to accumulate 100,000 hours of training data before 2027. This represents approximately 12 years of continuous real-time robotic operation compressed into several months of automated generation.

The synthesized datasets will serve as the primary training corpus for LG's proprietary Robot Foundation Model (RFM). Similar to how multimodal foundation models learn representations across text and images, the RFM architecture is designed to unify perception, environmental spatial reasoning, and motor actuation across diverse robotic form factors.

Hardware Integration and Group Restructuring

Alongside data pipelines, LG is exploring Nvidia's Isaac GR00T humanoid foundation model for modular robots and next-generation bipedal hardware. The companies plan to co-develop reference designs combining Nvidia compute hardware with LG's proprietary actuator mechanisms and manufacturing supply chains.

The initiative follows structural changes within LG. In July 2026, the company established a dedicated Robotics Business Center reporting directly to the CEO, consolidating robotic engineering across LG Electronics, LG Innotek, and LG CNS.

Sources

Written by

More to read

  • Fully Sharded Data Parallel (FSDP) and ZeRO: How Memory Sharding Eliminates Redundant Model States in Distributed Training

    Fully Sharded Data Parallel (FSDP) and ZeRO: How Memory Sharding Eliminates Redundant Model States in Distributed Training Training large language models across distributed GPU clusters introduces a fundamental memory bottleneck. In traditional Distributed Data Parallel (DDP) setups, every GPU maintains an identical copy of model weights, optimizer states, and gradients while processing independent data batches. As models scale from billions to hundreds of billions of parameters, static model s

    1 min
  • LLM Fine-Tuning Frameworks in Production: Unsloth vs. Axolotl vs. LLaMA-Factory vs. Torchtune Architecture, Throughput, and Distributed Scaling

    Modern post-training pipelines have moved beyond basic training scripts. As model parameter counts, context windows, and alignment techniques expand, the choice of fine-tuning framework directly dictates GPU memory overhead, token throughput, and developer iteration speed. Four open-source frameworks dominate the enterprise fine-tuning landscape: Unsloth, Axolotl, LLaMA-Factory, and Meta's Torchtune. While all four orchestrate parameter-efficient fine-tuning (PEFT) and full parameter adaptation

    1 min
  • Anthropic Prepares Dual-Class Super-Voting Shares for Co-Founders Ahead of Planned IPO

    Anthropic is preparing to implement a dual-class share structure that grants super-voting equity to its co-founders ahead of a planned initial public offering, according to a report from The Information. The mechanism is designed to concentrate long-term operational voting control with executive leadership and insulate decision-making from external market and investor pressures. The structure comes as the maker of the Claude model family scales enterprise commercialization, with annual revenue

    1 min