{"id":846,"date":"2025-12-04T15:09:31","date_gmt":"2025-12-04T07:09:31","guid":{"rendered":"https:\/\/www.midytech.com\/?p=846"},"modified":"2025-12-04T15:09:31","modified_gmt":"2025-12-04T07:09:31","slug":"this-model-will-revolutionize-how-humans-access-information","status":"publish","type":"post","link":"https:\/\/www.midytech.com\/index.php\/2025\/12\/04\/this-model-will-revolutionize-how-humans-access-information\/","title":{"rendered":"This Model Will Revolutionize How Humans Access Information"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">\u200bLast week, ByteDance released its latest model,\u00a0<strong>Vidi2<\/strong>, whose core capability is the rapid interpretation of videos \u2014 essentially analyzing every frame without human intervention and extracting corresponding data.This is the\u00a0<strong>VIDI2<\/strong>\u200b model.As a product manager, I\u2019ve always been keenly interested in groundbreaking technologies \u2014 especially those that, during my PhD research, I hoped would become engineering-driven products with strong technical barriers.<strong>A near-revolutionary technology: it changes how we access information.<\/strong>\u200bNowadays, turning WeChat Official Accounts posts into image carousels or videos is already a mainstream form of content creation. But what if we could reverse that process \u2014 turning videos back into text? This would dramatically improve the efficiency of content information flow and double humanity\u2019s ability to retrieve information.In the past, we used to say,\u00a0<em>\u201cWhere has someone been?\u201d<\/em>But now,\u00a0<strong>it\u2019s our ability to access and retrieve information that shapes each person\u2019s worldview.<\/strong>\u200bThis model will be nothing short of revolutionary for new media creators and influencers.Just like how I \u2014 and so many others \u2014 now primarily consume information through video, with short and long-form videos dominating the landscape, fewer and fewer people are reading text. As humans, we\u2019re naturally drawn to faster, higher-frequency consumption patterns \u2014 the so-called &#8220;lazy mode&#8221; of media engagement.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>Supports Keyword Search in Videos<\/strong>\u200b<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">With Vidi2, you can turn it into a powerful tool for new media \u2014 even for educational videos or robotic learning applications. It can extract the narrative and steps from a video in text form, then allow a large language model to compare and memorize the corresponding actions in the video, thereby accelerating model convergence.For example, in the official demo video:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\u2022You can search for scenes containing\u00a0<strong>a dragon<\/strong>\u200b and get a list of matching frames.<\/li>\n\n\n\n<li>\u2022Input \u201chand,\u201d and it will output all video segments where hands appear.<\/li>\n<\/ul>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>User-Acceptable Efficiency: From Text Search to Video Search<\/strong>\u200b<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">With the foundational technology of Vidi2, we can now move toward&nbsp;<strong>video-based search<\/strong>, rather than relying on titles.This means&nbsp;<strong>clickbait titles will become meaningless<\/strong>, and videos with deceptive thumbnails but unrelated content will no longer work.The focus will finally return to the&nbsp;<strong>actual content of the video<\/strong>\u200b \u2014 including any text within it that can be interpreted.Imagine the vast amount of content on the internet today. To truly find what you need by manually watching videos is incredibly time-consuming \u2014 especially when reviewing surveillance footage.But with this technology, you can&nbsp;<strong>search inside surveillance videos<\/strong>, quickly locate the exact clips you need, and save massive amounts of time.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>Supports Video Element Editing<\/strong>\u200b<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Beyond search, Vidi2 also enables&nbsp;<strong>editing of video elements<\/strong>. Users can search for specific objects and replace them, effectively transforming parts of the video into something else.It\u2019s reminiscent of a sci-fi movie \u2014 like&nbsp;<em>Bloodshot<\/em>starring Vin Diesel \u2014 where a tech company uses video editing to alter the protagonist\u2019s memory by manipulating objects, characters, and even dialogue in video-like reconstructed memories, eventually turning him into an assassin.The scene above shows&nbsp;<strong>memory editing<\/strong>, which is akin to spatial intelligence. While Vidi2 currently only supports&nbsp;<strong>2D video<\/strong>, not spatial or 3D video, it\u2019s already powerful enough to&nbsp;<strong>double the efficiency of how we access information today<\/strong>.The retrieval speed is now&nbsp;<strong>practically usable<\/strong>\u200b \u2014 far surpassing the experience of watching a short video, let alone sitting through an entire long-form one.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>\u200bLast week, ByteDance released&hellip;<\/p>\n","protected":false},"author":2,"featured_media":847,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[9,8],"tags":[24,254,26],"class_list":["post-846","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-companies","category-deep-tech","tag-ai","tag-bytedance","tag-deep-tech"],"_links":{"self":[{"href":"https:\/\/www.midytech.com\/index.php\/wp-json\/wp\/v2\/posts\/846","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.midytech.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.midytech.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.midytech.com\/index.php\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.midytech.com\/index.php\/wp-json\/wp\/v2\/comments?post=846"}],"version-history":[{"count":1,"href":"https:\/\/www.midytech.com\/index.php\/wp-json\/wp\/v2\/posts\/846\/revisions"}],"predecessor-version":[{"id":848,"href":"https:\/\/www.midytech.com\/index.php\/wp-json\/wp\/v2\/posts\/846\/revisions\/848"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.midytech.com\/index.php\/wp-json\/wp\/v2\/media\/847"}],"wp:attachment":[{"href":"https:\/\/www.midytech.com\/index.php\/wp-json\/wp\/v2\/media?parent=846"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.midytech.com\/index.php\/wp-json\/wp\/v2\/categories?post=846"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.midytech.com\/index.php\/wp-json\/wp\/v2\/tags?post=846"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}