{"id":147454,"date":"2025-01-31T19:00:00","date_gmt":"2025-01-31T19:00:00","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2025\/01\/31\/openai-launches-o3-mini-its-latest-reasoning-model-techcrunch\/"},"modified":"2025-01-31T19:00:00","modified_gmt":"2025-01-31T19:00:00","slug":"openai-launches-o3-mini-its-latest-reasoning-model-techcrunch","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2025\/01\/31\/openai-launches-o3-mini-its-latest-reasoning-model-techcrunch\/","title":{"rendered":"OpenAI launches o3-mini, its latest &#8216;reasoning&#8217; model | TechCrunch"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">OpenAI on Friday launched a new AI \u201creasoning\u201d model, o3-mini, the newest in the company\u2019s <a href=\"https:\/\/techcrunch.com\/tag\/o1\/\" target=\"_blank\" rel=\"noopener\">o family of reasoning models<\/a>.<\/p>\n<p class=\"wp-block-paragraph\">OpenAI <a href=\"https:\/\/techcrunch.com\/2024\/12\/20\/openai-announces-new-o3-model\/\" target=\"_blank\" rel=\"noopener\">first previewed the model in December<\/a> alongside a more capable system called o3, but the launch comes at a pivotal moment for the company, whose ambitions \u2014 and challenges \u2014 are seemingly growing by the day.<\/p>\n<p class=\"wp-block-paragraph\">OpenAI is battling the perception that it\u2019s ceding ground in the AI race to <a href=\"https:\/\/techcrunch.com\/2025\/01\/28\/deepseek-everything-you-need-to-know-about-the-ai-chatbot-app\/\" target=\"_blank\" rel=\"noopener\">Chinese companies like DeepSeek<\/a>, which OpenAI alleges might have stolen its IP. It has been trying to <a href=\"https:\/\/siliconangle.com\/2025\/01\/19\/openai-ceo-brief-us-officials-advanced-ai-agents-capable-complex-tasks\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">shore up its relationship with Washington<\/a> as it simultaneously pursues an <a href=\"https:\/\/techcrunch.com\/2025\/01\/21\/openai-teams-up-with-softbank-and-oracle-on-50b-data-center-project\/\" target=\"_blank\" rel=\"noopener\">ambitious data center project<\/a>, and <a href=\"https:\/\/techcrunch.com\/2025\/01\/30\/openai-said-to-be-in-talks-to-raise-40b-at-a-340b-valuation\/\" target=\"_blank\" rel=\"noopener\">as it reportedly lays the groundwork<\/a> for one of the largest funding rounds in history.<\/p>\n<p class=\"wp-block-paragraph\">Which brings us to o3-mini. OpenAI is pitching its new model as both \u201cpowerful\u201d and \u201caffordable.\u201d<\/p>\n<p class=\"wp-block-paragraph\">\u201cToday\u2019s launch marks [\u2026] an important step toward broadening accessibility to advanced AI in service of our mission,\u201d an OpenAI spokesperson told TechCrunch.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-more-efficient-reasoning\">More efficient reasoning<\/h2>\n<p class=\"wp-block-paragraph\">Unlike most large language models, reasoning models like o3-mini thoroughly fact-check themselves before giving out results. This\u00a0helps them <a href=\"https:\/\/techcrunch.com\/2024\/08\/27\/why-ai-cant-spell-strawberry\/\" target=\"_blank\" rel=\"noopener\">avoid some of the\u00a0pitfalls<\/a>\u00a0that normally trip up models. These reasoning models do take a little longer to arrive at solutions, but the trade-off is that they tend to be more reliable \u2014 though not perfect \u2014 in domains like physics.<\/p>\n<p class=\"wp-block-paragraph\">O3-mini is fine-tuned for STEM problems, specifically for programming, math, and science. OpenAI claims the model is largely on par with the o1 family, o1 and o1-mini, in terms of capabilities, but runs faster and costs less.<\/p>\n<p class=\"wp-block-paragraph\">The company claimed that external testers preferred o3-mini\u2019s answers over those from o1-mini more than half the time. O3-mini apparently also made 39% fewer \u201cmajor mistakes\u201d on \u201ctough real-world questions\u201d in <a href=\"https:\/\/hbr.org\/2017\/06\/a-refresher-on-ab-testing\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">A\/B tests<\/a> versus o1-mini, and produced \u201cclearer\u201d responses while delivering answers about 24% faster.<\/p>\n<p class=\"wp-block-paragraph\">O3-mini will be available to all users via <a href=\"https:\/\/techcrunch.com\/2025\/01\/28\/chatgpt-everything-to-know-about-the-ai-chatbot\/\" target=\"_blank\" rel=\"noopener\">ChatGPT<\/a> starting Friday, but users who pay for OpenAI\u2019s ChatGPT Plus and Team plans will get a higher rate limit of 150 queries per day. ChatGPT Pro subscribers will get unlimited access, and o3-mini will come to ChatGPT Enterprise and ChatGPT Edu customers in a week. (No word on <a href=\"https:\/\/techcrunch.com\/2025\/01\/28\/openai-launches-chatgpt-plan-for-u-s-government-agencies\/\" target=\"_blank\" rel=\"noopener\">ChatGPT Gov<\/a> yet).<\/p>\n<p class=\"wp-block-paragraph\">Users with premium plans can select o3-mini using the ChatGPT drop-down menu. Free users can click or tap the new \u201cReason\u201d button in the chat bar, or have ChatGPT \u201cre-generate\u201d an answer.<\/p>\n<p class=\"wp-block-paragraph\">Beginning Friday, o3-mini will also be available via OpenAI\u2019s API to select developers, but it initially will not have support for analyzing images. Devs can select the level of \u201creasoning effort\u201d (low, medium, or high) to get o3-mini to \u201cthink harder\u201d based on their use case and latency needs.<\/p>\n<p class=\"wp-block-paragraph\">O3-mini is priced at $0.55 per million cached input tokens and $4.40 per million output tokens, where a million tokens equates to roughly 750,000 words. That\u2019s 63% cheaper than o1-mini, and competitive with DeepSeek\u2019s R1 reasoning model pricing. DeepSeek charges $0.14 per million cached input tokens and $2.19 per million output tokens for R1 access through its API.<\/p>\n<p class=\"wp-block-paragraph\">In ChatGPT, o3-mini is set to medium reasoning effort, which OpenAI says provides \u201ca balanced trade-off between speed and accuracy.\u201d Paid users will have the option of selecting \u201co3-mini-high\u201d in the model picker, which will deliver what OpenAI calls \u201chigher intelligence\u201d in exchange for slower responses.<\/p>\n<p class=\"wp-block-paragraph\">Regardless of which version of o3-mini ChatGPT users choose, the model will work with search to find up-to-date answers with links to relevant web sources. OpenAI cautions that the functionality is a \u201cprototype\u201d as it works to integrate search across its reasoning models.<\/p>\n<p class=\"wp-block-paragraph\">\u201cWhile o1 remains our broader general-knowledge reasoning model, o3-mini provides a specialized alternative for technical domains requiring precision and speed,\u201d OpenAI wrote in a blog post on Friday. \u201cThe release of o3-mini marks another step in OpenAI\u2019s mission to push the boundaries of cost-effective intelligence.\u201d<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-caveats-abound\">Caveats abound<\/h2>\n<p class=\"wp-block-paragraph\">O3-mini is not OpenAI\u2019s most powerful model to date, nor does it leapfrog DeepSeek\u2019s R1 reasoning model in every benchmark.<\/p>\n<p class=\"wp-block-paragraph\">O3-mini beats R1 on AIME 2024, a test\u00a0that measures how well models understand and respond to complex instructions \u2014 but only with high reasoning effort. It also beats R1 on the programming-focused test SWE-bench Verified (by .1 point), but again, only with high reasoning effort. On low reasoning effort, o3-mini lags R1 on GPQA Diamond, which tests models with PhD-level physics, biology, and chemistry questions.<\/p>\n<p class=\"wp-block-paragraph\">To be fair, o3-mini answers many queries at competitively low cost and latency. In the post, OpenAI compares its performance to the o1 family:<\/p>\n<p class=\"wp-block-paragraph\">\u201cWith low reasoning effort, o3-mini achieves comparable performance with o1-mini, while with medium effort, o3-mini achieves comparable performance with o1,\u201d OpenAI writes. \u201cO3-mini with medium reasoning effort matches o1\u2019s performance in math, coding and science while delivering faster responses. Meanwhile, with high reasoning effort, o3-mini outperforms both o1-mini and o1.\u201d<\/p>\n<p class=\"wp-block-paragraph\">It\u2019s worth noting that o3-mini\u2019s performance advantage over o1 is slim in some areas. On AIME 2024, o3-mini beats o1 by just 0.3 percentage points when set to high reasoning effort. And on GPQA Diamond, o3-mini doesn\u2019t surpass o1\u2019s score even on high reasoning effort.<\/p>\n<p class=\"wp-block-paragraph\">OpenAI asserts that o3-mini is as \u201csafe\u201d or safer than the o1 family, however, thanks to red-teaming efforts and its \u201cdeliberative alignment\u201d methodology, which makes models \u201cthink\u201d about OpenAI\u2019s safety policy while they\u2019re responding to queries. According to the company, o3-mini \u201csignificantly surpasses\u201d one of OpenAI\u2019s flagship models, <a href=\"https:\/\/techcrunch.com\/2024\/05\/13\/openais-newest-model-is-gpt-4o\/\" target=\"_blank\" rel=\"noopener\">GPT-4o<\/a>, on \u201cchallenging safety and jailbreak evaluations.\u201d<\/p>\n<p class=\"wp-block-paragraph\"><em>TechCrunch has an AI-focused newsletter!\u00a0<a href=\"https:\/\/techcrunch.com\/newsletters\/\" target=\"_blank\" rel=\"noreferrer noopener\">Sign up here<\/a>\u00a0to get it in your inbox every Wednesday.<\/em><\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/techcrunch.com\/2025\/01\/31\/openai-launches-o3-mini-its-latest-reasoning-model\/\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>OpenAI on Friday launched a new AI \u201creasoning\u201d model, o3-mini, the newest in the company\u2019s o family of reasoning models. OpenAI first previewed the model in December alongside a more capable system called o3, but the launch comes at a pivotal moment for the company, whose ambitions \u2014 and challenges \u2014 are seemingly growing by [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":147455,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-147454","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/147454","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=147454"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/147454\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/147455"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=147454"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=147454"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=147454"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}