Rednote Releases First Open-Weight Model in the dots3 Series, Exploring Long-Horizon Real-World Tasks

by shayaan

SHANGHAI, CHINA / ACCESS Newswi e / August 16, 2026 / O August 14, dots studio, ed ote’s model lab, eleased the model weights fo dots3 ote p eview. The model belo gs to the same dots3 se ies as the model that p eviously achieved a pe fect sco e of 42 poi ts at the 2026 I te atio al Mathematical Olympiad (IMO), a d ma ks the fi st ope -weight elease i the dots3 se ies.

The model has 280B total pa amete s a d 16B active pa amete s, suppo ts a co text wi dow of up to 512K toke s, a d offe s multimodal u de sta di g ac oss text, visio , a d audio. It has also bee optimized fo complex easo i g a d lo g-ho izo age t tasks. Ac oss a a ge of mai st eam be chma k esults, dots3 – ote p eview a ks amo g the leadi g Chi ese models of compa able size i easo i g, age t capabilities, a d multimodal pe ceptio , with pa ticula ly st o g visual capabilities amo g models of a simila size.


(St o ge p oblem-solvi g capabilities)


(Multimodal capabilities)

Figu e: Evaluatio esults of dots3- ote-p eview o mai st eam Be chma k

The team believes that lo g-ho izo eal-wo ld tasks will become a impo ta t ext f o tie fo adva ci g la ge-model capabilities, yet the i dust y has ot paid e ough atte tio to this a ea. The team p eviously i t oduced two be chma ks built a ou d complex tasks i eve yday-life sce a ios. Leadi g models wo ldwide pe fo med poo ly o these evaluatio s, with o e eachi g the passi g th eshold. Th ough the ope -weight elease of dots3 ote, the team hopes to sha e its latest tech ical thi ki g with the b oade i dust y a d e cou age fu the explo atio of this di ectio .

Ove the past few yea s, la ge models have adva ced apidly o elatively closed-e ded tasks such as mathematics, codi g, a d e gi ee i g. These tasks sha e a impo ta t cha acte istic: it is elatively easy to dete mi e whethe a a swe is ight o w o g, a d models ca eceive clea feedback o how well they pe fo m, maki g them easie to t ai a d co ti uously imp ove.

See also  Somantra AI Releases AI Search Ranking Factors Report for Australian Insurance Brands, Analysing 2.4 Million ChatGPT and Google Citations

Real-wo ld tasks, howeve , a e fa mo e complex. Pla i g a t ip, e ovati g a home, o o ga izi g a weddi g may u fold ove days o eve mo ths. Use s may ot be able to a ticulate all of thei co st ai ts a d p efe e ces at the outset, while exte al facto s such as p ice fluctuatio s, flight cha ges, a d cha gi g weathe co ditio s may eme ge alo g the way.

As a esult, models eed ot o ly to u de sta d i fo matio beyo d text, i cludi g images a d audio, but also to co ti uously assess whethe thei cu e t pla s a e effective th oughout lo g-ho izo tasks a d imp ove them alo g the way, athe tha waiti g u til the e d to dete mi e success o failu e.

dots studio focuses o imp ovi g la ge models’ pe fo ma ce o lo g-ho izo eal-wo ld tasks. This alig s with the studio’s missio to “C eate f o tie i tellige ce fo daily life,” while co ti ui g ed ote’s lo gsta di g focus o eve ydaylife sce a ios a d its belief i usi g tech ology to be efit o di a y people.

The ewly eleased dots3 – ote p eview ep ese ts the latest p og ess i dots studio’s effo ts to build age ts capable of ha dli g lo g-ho izo eal-wo ld tasks. Ac oss multiple easo i g a d age t tasks, the model ca match o eve outpe fo m much la ge models with seve al times its pa amete cou t.

To add ess the challe ges of eal-wo ld tasks, the tech ical app oach behi d dots3 – ote p eview focuses o th ee a eas: Fi st, st o ge multimodal u de sta di g. I fo matio i eve yday-life sce a ios is ot limited to text. Floo pla s, quotatio s, flight sc ee shots, maps, a d voice memos all co tai complex fo ms of i fo matio that a model eeds to u de sta d befo e it ca effectively ha dle eal-wo ld tasks.

See also  The Future of NFTs: What Comes After the Hype (2025–2030 Outlook)

Seco d, the i t oductio of self-c itiqui g. Lo g-ho izo tasks ca ot ely solely o spa se ewa ds at the e d of a t ajecto y. Du i g RL t ai i g, the team t ai s the model ot o ly to solve p oblems, but also to evaluate its ow p og ess alo g the way. By lea i g to c itique i te mediate states-ide tifyi g mistakes, eassessi g its hypotheses, a d estimati g whethe it is o the ight t ack-the model eceives iche lea i g sig als th oughout lo g t ajecto ies, e abli g a mo e scalable ei fo ceme t lea i g pa adigm.

The same self-c itiqui g capability ca also be scaled at i fe e ce time. Rathe tha elyi g o a si gle attempt, the model ca ite atively eview a d efi e its ow solutio s. At IMO 2026, this app oach demo st ated its pote tial: th ough ite ative self-c itiqui g, the dots3 ote se ies achieved a officially ce tified pe fect sco e of 42/42 a d a gold medal.

Thi d, TEMPO a d lo g-ho izo ei fo ceme t lea i g. Fo tasks that last fo weeks o eve lo ge , co ve tio al ei fo ceme t lea i g st uggles to accu ately att ibute a fi al outcome to each i te mediate decisio : task t ajecto ies a e too lo g, while mea i gful feedback is too spa se. To add ess this bottle eck, the team developed TEMPO, which demo st ated st o ge pe fo ma ce tha GRPO o ARC-AGI 3.

Real-wo ld tasks a e difficult to measu e usi g existi g be chma ks. To add ess this, the team developed two evaluatio f amewo ks.

O e of them, VibeSea chBe ch, p ima ily evaluates a model’s multi-tu sea ch capabilities as use s’ eeds g adually become clea e . It cove s 20 domai s a d 200 tasks, simulati g diffe e t use pe so as a d equi i g models to p og essively ide tify a d fill i missi g equi eme ts th ough multiple ou ds of i te actio .

The othe evaluatio f amewo k, VibeLifeBe ch, focuses o whethe a model ca follow th ough o tasks ove exte ded pe iods i a co ti uously cha gi g e vi o me t. It simulates the passage of eal time as well as exte al cha ges such as p ices a d se vice status. The evaluatio cove s 10 domai s a d 20 lo g-ho izo tasks, with each task spa i g 20 to 30 stages a d a total of 1,247 evaluatio checks.

See also  VanEck bets BNB’s real-world usage can stand out in a crowded crypto ETF market

Cu e t esults suggest that leadi g models still have substa tial oom fo imp oveme t o both evaluatio s. O VibeLifeBe ch, fo example, all seve leadi g models tested fell below the passi g th eshold, with Claude Opus 5 a ki g fi st at 0.325. O VibeSea chBe ch, Claude Opus 5 led with a sco e of 31.14, while GPT-5.4 a ked last.

These esults also suggest that, as models move f om ve ifiable tasks towa d eal-wo ld tasks, thei capabilities still face sig ifica t challe ges.

The full ve sio of dots3 ote is also expected to be eleased with ope weights i the ea futu e. dots3 ote is the lightest ve sio i the dots 3 se ies, a d the complete dots3 se ies will i clude th ee tie s- ote, jazz, a d a ia-desig ed fo applicatio s with diffe e t equi eme ts fo task complexity, espo se speed, a d compute cost. dots has p eviously eleased the model weights fo the text la ge la guage model dots.llm1, the multili gual docume t layout pa si g model dots.oc , a d the multimodal visual u de sta di g la ge model dots.vlm1.

Tech blog: https://studio.dots.ai/dots/dots3-e .htmlHuggi gFace: https://huggi gface.co/dots-studio/dots3- ote-p evdots studio website: https://studio.dots.ai/?la g=e

Compa y: dots studio ( ed ote/Xiaoho gshu)Co tact: Chao QiaoEmail: [email p otected]Website: https://studio.dots.ai/?la g=e

SOURCE: Dots Studio ( ed ote/Xiaoho gshu)

About Web3Wire
Web3Wire – Information, news, press releases, events and research articles about Web3, Metaverse, Blockchain, Artificial Intelligence, Cryptocurrencies, Decentralized Finance, NFTs and Gaming.
Visit Web3Wire for Web3 News and Events, Block3Wire for the latest Blockchain news and Meta3Wire to stay updated with Metaverse News.

web3wire.org

You may also like

Latest News

Copyright © Sovereign Wealth Signals