New Deepseek, human genome map, Navier Stokes, GPT finance, Suno v6, YuE2: AI NEWS
Summary
This video showcases a rapid succession of AI advancements across various domains. Key developments include new state-of-the-art open-source models for image analysis (Merryold V2), character animation (Unimate), music generation (UA2), and large language models (Deepseek 4.1 Flash), along with significant progress in robotics (Isaac 0.5, Uni LMWLA) and 3D reconstruction (World Sculpt, Fire 3D). Notably, OpenAI utilized a swarm of agents to solve a long-standing math problem (Navier-Stokes), and Google DeepMind created a comprehensive map of human genome mutations. The video also highlights practical applications like advanced speech editing (Out), real-time world generation (Lingbot World 2), and efficient large model execution on limited hardware (Edge Zero).
Key Insights
Merryold V2 enhances image analysis with pixel-level 3D depth and surface normal estimation.
Merryold V2 interprets regular images to generate detailed depth maps, surface normals, and albedo maps with pixel-level accuracy, outperforming previous models like Moji 3 and Infinidepth in detail and resolution.
Unimate animates custom 3D skeletons across diverse and unusual character types.
Unimate utilizes a single AI model to animate a wide variety of 3D skeleton types, including non-human characters like flowers, dragons, and satellites, by taking a rigged 3D model and text instructions, without requiring further retraining.
Google DeepMind's Alpha Genome Atlas maps effects of billions of human genome mutations.
Alpha Genome Atlas uses AI to predict the effects of approximately 9 billion single-letter mutations in the human genome, creating a massive dataset that allows scientists to quickly look up mutation impacts and identify significant variants.
Isaac 0.5 provides a general foundation model for robot perception and action prediction.
Isaac 0.5 is a 36 billion parameter sparse model trained on diverse robot data and general video, enabling robots to understand visual input, predict future states, and generate movements by integrating various sensor data.
Tencent's Out offers advanced speech synthesis, cloning, and micro-editing capabilities.
Out provides text-to-speech, zero-shot voice cloning, seamless micro-editing of existing speech (adding/deleting words), noise reduction, quality enhancement, emotion change, and even timbre alteration.
UA2 is a high-quality open-source music generator that composes musical plans before generation.
UA2 first creates a structured musical plan (melody, rhythm, chords) from a prompt, allowing for detailed editing before generating the final song, claimining to surpass even Sunno V6 in quality.
Deepseek 4.1 Flash sets new benchmarks for open-source LLMs, rivaling top proprietary models.
This MoE model, surprisingly efficient with only 8-16B active parameters, achieves state-of-the-art on multiple benchmarks, including Deep Suite 1.1 and Automation Bench, while offering impressive speed and cost efficiency via API.
OpenAI's AI swarm solves the 90-year-old Navier-Stokes Millennium Prize problem.
Using 10,000 agents over 88 hours, OpenAI's internal model proved a singularity exists in the forced, non-Eulerian Navier-Stokes equations, a feat unsolved by humans for decades.
Edge Zero enables running large Mixture-of-Experts LLMs on low-memory devices.
By streaming only necessary experts, Edge Zero allows models like Quen 3.5B to run with significantly reduced memory, down to 2.9 GB, preserving quality through techniques like RecoverLoRA.
Uni LMWLA is a compact 6B parameter robot model capable of diverse tasks and whole-body coordination.
This model integrates vision, instructions, and state to predict scene changes and generate actions for 64 tasks, supporting various grippers and demonstrating capabilities from laundry loading to trash disposal.
Sections
3D Image Understanding and Animation
Merryold V2 enhances image analysis with pixel-level 3D depth and surface normal estimation.
Merryold V2 interprets regular images to generate detailed depth maps, surface normals, and albedo maps with pixel-level accuracy, outperforming previous models like Moji 3 and Infinidepth in detail and resolution.
Unimate animates custom 3D skeletons across diverse and unusual character types.
Unimate utilizes a single AI model to animate a wide variety of 3D skeleton types, including non-human characters like flowers, dragons, and satellites, by taking a rigged 3D model and text instructions, without requiring further retraining.
Genomics and Interactive Worlds
Google DeepMind's Alpha Genome Atlas maps effects of billions of human genome mutations.
Alpha Genome Atlas uses AI to predict the effects of approximately 9 billion single-letter mutations in the human genome, creating a massive dataset that allows scientists to quickly look up mutation impacts and identify significant variants.
Lingbot World 2 generates high-quality, real-time interactive virtual environments.
Lingbot World 2 continuously generates interactive virtual worlds controllable via key presses and text prompts. It achieves 720p at 60fps, supports NPCs, and streams worlds chunk by chunk for real-time generation.
Robotics and 3D Scene Reconstruction
Isaac 0.5 provides a general foundation model for robot perception and action prediction.
Isaac 0.5 is a 36 billion parameter sparse model trained on diverse robot data and general video, enabling robots to understand visual input, predict future states, and generate movements by integrating various sensor data.
World Sculpt reconstructs scenes into editable, individual 3D objects.
World Sculpt converts images or videos into 3D scenes with distinct, editable objects, differentiating from other systems by separating and positioning each element correctly within the shared world.
Fire 3D rapidly creates simulation-ready 3D scenes with editable, meshed objects.
Fire 3D reconstructs 3D scenes from photos or videos in under a minute, generating simulation-ready scenes with individual meshes and material information for each object, processing up to 16 objects in parallel.
Advanced Speech and Music Generation
Tencent's Out offers advanced speech synthesis, cloning, and micro-editing capabilities.
Out provides text-to-speech, zero-shot voice cloning, seamless micro-editing of existing speech (adding/deleting words), noise reduction, quality enhancement, emotion change, and even timbre alteration.
UA2 is a high-quality open-source music generator that composes musical plans before generation.
UA2 first creates a structured musical plan (melody, rhythm, chords) from a prompt, allowing for detailed editing before generating the final song, claimining to surpass even Sunno V6 in quality.
Large Language Models and Benchmarking
Deepseek 4.1 Flash sets new benchmarks for open-source LLMs, rivaling top proprietary models.
This MoE model, surprisingly efficient with only 8-16B active parameters, achieves state-of-the-art on multiple benchmarks, including Deep Suite 1.1 and Automation Bench, while offering impressive speed and cost efficiency via API.
Real Suite benchmark tests AI's ability to perform practical, business-critical software engineering tasks.
Unlike previous benchmarks, Real Suite evaluates AI agents on real-world tasks involving private company code and business logic, where even top models like Fable 5.1 and GPT-4 score below 40%.
AI in Scientific Discovery and Robotics Control
OpenAI's AI swarm solves the 90-year-old Navier-Stokes Millennium Prize problem.
Using 10,000 agents over 88 hours, OpenAI's internal model proved a singularity exists in the forced, non-Eulerian Navier-Stokes equations, a feat unsolved by humans for decades.
Show Harness enables control of robots using any vision-language model.
This framework allows vision-language models like Gemini or GPT-4 to interpret scenes and generate robot-specific actions, achieving up to 100% success rate in controlling robot movements.
Edge Zero enables running large Mixture-of-Experts LLMs on low-memory devices.
By streaming only necessary experts, Edge Zero allows models like Quen 3.5B to run with significantly reduced memory, down to 2.9 GB, preserving quality through techniques like RecoverLoRA.
UMR translates human motion to robot movements, respecting robot constraints.
Unified Motion Retargeting represents bodies as point clouds to map human movements onto humanoid robots, maintaining realism and ensuring robots can perform complex actions like grasping or climbing.
Uni LMWLA is a compact 6B parameter robot model capable of diverse tasks and whole-body coordination.
This model integrates vision, instructions, and state to predict scene changes and generate actions for 64 tasks, supporting various grippers and demonstrating capabilities from laundry loading to trash disposal.
Specialized AI Tools and Models
OpenBM's MiniCPM 2B offers state-of-the-art performance for tiny, offline edge devices.
This 2 billion parameter dense model outperforms other small models across various benchmarks and is released with its high-quality training dataset, suitable for devices with limited resources.
Ask a Question
*Uses 1 Wisdom coin from your coin balance











