{"id":89786,"date":"2024-04-13T16:15:41","date_gmt":"2024-04-13T16:15:41","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2024\/04\/13\/vana-plans-to-let-users-rent-out-their-reddit-data-to-train-ai-techcrunch\/"},"modified":"2024-04-13T16:15:41","modified_gmt":"2024-04-13T16:15:41","slug":"vana-plans-to-let-users-rent-out-their-reddit-data-to-train-ai-techcrunch","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2024\/04\/13\/vana-plans-to-let-users-rent-out-their-reddit-data-to-train-ai-techcrunch\/","title":{"rendered":"Vana plans to let users rent out their Reddit data to train AI | TechCrunch"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p id=\"speakable-summary\"><span class=\"featured__span-first-words\">In the generative<\/span> AI boom, data is the new oil. So why shouldn\u2019t you be able to sell your own?<\/p>\n<p>From big tech firms to startups, AI makers are licensing e-books, images, videos, audio and more from data brokers, all in the pursuit of training up more capable (<a href=\"https:\/\/techcrunch.com\/2023\/01\/27\/the-current-legal-cases-against-generative-ai-are-just-the-beginning\/\" target=\"_blank\" rel=\"noopener\">and more legally defensible<\/a>) AI-powered products. Shutterstock has <a href=\"https:\/\/techcrunch.com\/2023\/07\/11\/shutterstock-expands-deal-with-openai-to-build-generative-ai-tools\/\" target=\"_blank\" rel=\"noopener\">deals<\/a> with Meta, Google, Amazon and Apple to supply millions of images for model training, while OpenAI has <a href=\"https:\/\/techcrunch.com\/2024\/03\/13\/are-openais-deals-with-publishers-edging-out-the-competition\/#:~:text=How%20much%20is%20OpenAI%20paying,to%20train%20its%20GenAI%20models.\" target=\"_blank\" rel=\"noopener\">signed agreements<\/a> with several news organizations to train its models on news archives.<\/p>\n<p>In many cases, the individual creators and owners of that data haven\u2019t seen a dime of the cash changing hands. A startup called <a href=\"http:\/\/vana.com\" target=\"_blank\" rel=\"noopener\">Vana<\/a> wants to change that.<\/p>\n<p>Anna Kazlauskas and Art Abal, who met in a class at the MIT Media Lab focused on building tech for emerging markets, co-founded Vana in 2021. Prior to Vana, Kazlauskas studied computer science and economics at MIT, eventually leaving to launch a fintech automation startup, Iambiq, out of Y Combinator. Abal, a corporate lawyer by training and education, was an associate at The Cadmus Group, a Boston-based consulting firm, before heading up impact sourcing at data annotation company Appen.<\/p>\n<p>With Vana, Kazlauskas and Abal set out to build a platform that lets users \u201cpool\u201d their data \u2014 including chats, speech recordings and photos \u2014 into data sets that can then be used for generative AI model training. They also want to create more personalized experiences \u2014 for instance, daily motivational voicemail based on your wellness goals, or an art-generating app that understands your style preferences\u00a0 \u2014 by fine-tuning public models on that data.<\/p>\n<p>\u201cVana\u2019s infrastructure in effect creates a user-owned data treasury,\u201d Kazlauskas told TechCrunch. \u201cIt does this by allowing users to aggregate their personal data in a non-custodial way \u2026 Vana allows users to own AI models and use their data across AI applications.\u201d<\/p>\n<p>Here\u2019s how Vana <a href=\"https:\/\/docs.vana.com\/api\" target=\"_blank\" rel=\"noopener\">pitches its platform and API to developers<\/a>:<\/p>\n<blockquote>\n<p>The Vana API connects a user\u2019s cross-platform personal data \u2026 to allow you to personalize your application. Your app gains instant access to a user\u2019s personalized AI model or underlying data, simplifying onboarding and eliminating compute cost concerns \u2026 We think users should be able to bring their personal data from walled gardens, like Instagram, Facebook and Google, to your application, so you can create amazing personalized experience from the very first time a user interacts with your consumer AI application.<\/p>\n<\/blockquote>\n<p>Creating an account with Vana is fairly simple. After confirming your email, you can attach data to a digital avatar (like selfies, a description of yourself and voice recordings) and explore apps built using Vana\u2019s platform and data sets. The app selection ranges from ChatGPT-style chatbots and interactive storybooks to a Hinge profile generator.<\/p>\n<div id=\"attachment_2691312\" style=\"width: 444px\" class=\"wp-caption aligncenter\"><\/p>\n<p id=\"caption-attachment-2691312\" class=\"wp-caption-text\"><strong>Image Credits:<\/strong> Vana<\/p>\n<\/div>\n<p>Now why, you might ask \u2014 in this age of increased data privacy awareness and ransomware attacks \u2014 would someone ever volunteer their personal info to an anonymous startup, much less a venture-backed one? (Vana has raised $20 million to date from Paradigm, Polychain Capital and other backers.) Can any profit-driven company really be trusted not to abuse or mishandle any monetizable data it gets its hands on?<\/p>\n<div id=\"attachment_2691315\" style=\"width: 475px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-2691315\" class=\"vertical size-full wp-image-2691315\" src=\"https:\/\/techcrunch.com\/wp-content\/uploads\/2024\/04\/Screenshot-2024-04-12-at-3.15.54\u202fPM.png\" alt=\"Vana Reddit DAO\" width=\"465\" height=\"984\" srcset=\"https:\/\/techcrunch.com\/wp-content\/uploads\/2024\/04\/Screenshot-2024-04-12-at-3.15.54\u202fPM.png 465w, https:\/\/techcrunch.com\/wp-content\/uploads\/2024\/04\/Screenshot-2024-04-12-at-3.15.54\u202fPM.png?resize=71,150 71w, https:\/\/techcrunch.com\/wp-content\/uploads\/2024\/04\/Screenshot-2024-04-12-at-3.15.54\u202fPM.png?resize=142,300 142w, https:\/\/techcrunch.com\/wp-content\/uploads\/2024\/04\/Screenshot-2024-04-12-at-3.15.54\u202fPM.png?resize=321,680 321w, https:\/\/techcrunch.com\/wp-content\/uploads\/2024\/04\/Screenshot-2024-04-12-at-3.15.54\u202fPM.png?resize=24,50 24w\" sizes=\"auto, (max-width: 465px) 100vw, 465px\"\/><\/p>\n<p id=\"caption-attachment-2691315\" class=\"wp-caption-text\"><strong>Image Credits:<\/strong> Vana<\/p>\n<\/div>\n<p>In response to that question, Kazlauskas stressed that the whole point of Vana is for users to \u201creclaim control over their data,\u201d noting that Vana users have the option to self-host their data rather than store it on Vana\u2019s servers and control how their data\u2019s shared with apps and developers. She also argued that, because Vana makes money by charging users a monthly subscription (starting at $3.99) and levying a \u201cdata transaction\u201d fee on devs (e.g. for transferring data sets for AI model training), the company is disincentivized to exploit users and the troves of personal data they bring with them.<\/p>\n<p>\u201cWe want to create models owned and governed users who all contribute their data,\u201d Kazlauskas said, \u201cand allow users to bring their data and models with them to any application.\u201d<\/p>\n<p>Now, while <em>Vana <\/em>isn\u2019t selling users\u2019 data to companies for generative AI model training (or so it claims), it wants to allow users to do this themselves if they choose \u2014 starting with their Reddit posts.<\/p>\n<p>This month, Vana launched what it\u2019s calling the <a href=\"https:\/\/www.rdatadao.org\/\" target=\"_blank\" rel=\"noopener\">Reddit Data DAO (Digital Autonomous Organization)<\/a>, a program that pools multiple users\u2019 Reddit data (including their karma and post history) and lets them to decide together how that combined data is used. After joining with a Reddit account, submitting a <a href=\"https:\/\/www.reddit.com\/settings\/data-request\" target=\"_blank\" rel=\"noopener\">request<\/a> to Reddit for their data and uploading that data to the DAO, users gain the right to vote alongside other members of the DAO on decisions like licensing the combined data to generative AI companies for a shared profit.<\/p>\n<div class=\"embed breakout embed-oembed embed--twitter\">\n<blockquote class=\"twitter-tweet\" data-width=\"550\" data-dnt=\"true\">\n<p lang=\"en\" dir=\"ltr\">We have crunched the numbers and r\/datadao is now largest data DAO in history: Phase 1 welcomed 141,000 reddit users with 21,000 full data uploads.<\/p>\n<p>\u2014 r\/datadao (@rdatadao) <a href=\"https:\/\/twitter.com\/rdatadao\/status\/1778263621631611231?ref_src=twsrc%5Etfw\" target=\"_blank\" rel=\"noopener\">April 11, 2024<\/a><\/p>\n<\/blockquote>\n<\/div>\n<p>It\u2019s an answer of sorts to Reddit\u2019s <a href=\"https:\/\/techcrunch.com\/2024\/02\/22\/reddit-says-its-made-203m-so-far-licensing-its-data\/\" target=\"_blank\" rel=\"noopener\">recent moves<\/a> to commercialize data on its platform.<\/p>\n<p>Reddit previously didn\u2019t gate access to posts and communities for generative AI training purposes. But it reversed course late last year, ahead of its IPO. Since the policy change, Reddit has raked in over $203 million in licensing fees from companies including Google.<\/p>\n<p>\u201cThe broad idea [with the DAO is] to free user data from the major platforms that seek to hoard and monetize it,\u201d Kazlauskas said. \u201cThis is a first and is part of our push to help people pool their data into user-owned data sets for training AI models.\u201d<\/p>\n<p>Unsurprisingly, Reddit \u2014 which isn\u2019t working with Vana in any official capacity \u2014 isn\u2019t pleased about the DAO.<\/p>\n<p>Reddit banned Vana\u2019s <a href=\"https:\/\/www.reddit.com\/r\/datadao\/\" target=\"_blank\" rel=\"noopener\">subreddit<\/a> dedicated to discussion about the DAO. And a Reddit spokesperson accused Vana of \u201cexploiting\u201d its data export system, which is designed to comply with data privacy regulations like the GDPR and California Consumer Privacy Act.<\/p>\n<p>\u201cOur data arrangements allow us to put guardrails on such entities, even on public information,\u201d the spokesperson told TechCrunch. \u201cReddit does not share non-public, personal data with commercial enterprises, and when Redditors request an export of their data from us, they receive non-public personal data back from us in accordance with applicable laws. Direct partnerships between Reddit and vetted organizations, with clear terms and accountability, matters, and these partnerships and agreements prevent misuse and abuse of people\u2019s data.\u201d<\/p>\n<p>But does Reddit have any real reason to be concerned?<\/p>\n<p>Kazlauskas envisions the DAO growing to the point where it impacts the amount Reddit can charge customers for its data. That\u2019s a long ways off, assuming it ever happens; the DAO has just over 141,000 members, a tiny fraction of Reddit\u2019s 73-million-strong user base. And some of those members could be bots or duplicate accounts.<\/p>\n<p>Then there\u2019s the matter of how to fairly distribute payments that the DAO might receive from data buyers.<\/p>\n<p>Currently, the DAO awards \u201ctokens\u201d \u2014 cryptocurrency \u2014 to users corresponding to their Reddit <a href=\"https:\/\/www.reddit.com\/r\/NewToReddit\/comments\/18ds3h3\/new_to_reddit_how_does_karma_work\/\" target=\"_blank\" rel=\"noopener\">karma<\/a>. But karma might not be the best measure of quality contributions to the data set \u2014 particularly in smaller Reddit communities with fewer opportunities to earn it.<\/p>\n<p>Kazlauskas floats the idea that members of the DAO could choose to share their cross-platform and demographic data, making the DAO potentially more valuable and incentivizing sign-ups. But that would also require users to place even more trust in Vana to treat their sensitive data responsibly.<\/p>\n<p>Personally, I don\u2019t see Vana\u2019s DAO reaching critical mass. The roadblocks standing in the way are far too many. I do think, however, that it won\u2019t be the last grassroots attempt to assert control over the data increasingly being used to train generative AI models.<\/p>\n<p>Startups like <a href=\"https:\/\/techcrunch.com\/2023\/05\/03\/spawning-lays-out-its-plans-for-letting-creators-opt-out-of-generative-ai-training\/\" target=\"_blank\" rel=\"noopener\">Spawning<\/a> are working on ways to allow creators to impose rules guiding how their data is used for training while vendors like Getty Images, Shutterstock and Adobe continue to <a href=\"https:\/\/techcrunch.com\/2023\/09\/30\/how-much-can-artists-make-from-generative-ai-vendors-wont-say\/\" target=\"_blank\" rel=\"noopener\">experiment with compensation schemes<\/a>. But no one\u2019s cracked the code yet. Can it even <em>be<\/em> cracked? Given the <a href=\"https:\/\/www.theverge.com\/2024\/4\/6\/24122915\/openai-youtube-transcripts-gpt-4-training-data-google\" target=\"_blank\" rel=\"noopener\">cutthroat<\/a> <a href=\"https:\/\/www.bloomberg.com\/news\/articles\/2024-04-04\/youtube-says-openai-training-sora-with-its-videos-would-break-the-rules#:~:text=YouTube%20Says%20OpenAI%20Training%20Sora%20With%20Its%20Videos%20Would%20Break%20Rules,-Neal%20Mohan%2C%20chief&amp;text=The%20use%20of%20YouTube%20videos,Executive%20Officer%20Neal%20Mohan%20said.\" target=\"_blank\" rel=\"noopener\">nature<\/a> of the generative AI industry, it\u2019s certainly a tall order. But perhaps someone will find a way \u2014 or policymakers will force one.<\/p>\n<\/p><\/div>\n<p><script async src=\"\/\/platform.twitter.com\/widgets.js\" charset=\"utf-8\"><\/script><br \/>\n<br \/><br \/>\n<br \/><a href=\"https:\/\/techcrunch.com\/2024\/04\/13\/vana-plans-to-let-users-rent-out-their-reddit-data-to-train-ai\/\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>In the generative AI boom, data is the new oil. So why shouldn\u2019t you be able to sell your own? From big tech firms to startups, AI makers are licensing e-books, images, videos, audio and more from data brokers, all in the pursuit of training up more capable (and more legally defensible) AI-powered products. Shutterstock [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":89787,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-89786","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/89786","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=89786"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/89786\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/89787"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=89786"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=89786"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=89786"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}