U.S. bombs Iranian nuclear sites (bbc.co.uk)

There is a useful feature in DeepSeek that isn't present in other commercial LLMs. It displays its internal "thinking" process. I wonder what technological aspect makes this possible. Do several LLMs communicate with each other before providing a solution? Are there different roles within these LLMs, such as some proposing solutions, others contradicting or offering alternative viewpoints, or reminding of overlooked aspects?

Comments (2)

123yawaworht456 · 8h ago

>Do several LLMs communicate with each other before providing a solution?

>I wonder what technological aspect makes this possible.

one of its training datasets (prioritized somehow over the rest of them) contains a large number of examples emulating the thinking process within <think></think> tags before providing an output. the model then emulates it at runtime.

JPLeRouzic · 2h ago

Thank you for taking the time to answer. However I am not sure the answer is "NO" because DeepSeek has a particular technique in their architecture. To cite this blog [0]:

"Modern large language models (LLMs) started introducing a layer called “Mixture of Experts” (MoE) in their Transformer blocks to scale parameter count without linearly increasing compute. This is typically done through top-k (often k=2) “expert routing”, where each token is dispatched to two specialized feed-forward networks (experts) out of a large pool.

A naive GPU cluster implementation would be to place each expert on a separate device and have the router dispatch to the selected experts during inference. But this would have all the non-active experts idle on the expensive GPUs.

GShard, 2021 introduced the concept of sharding these feed-forward (FF) experts across multiple devices, so that each device"

[0] https://www.kernyan.com/hpc,/cuda/2025/02/26/Deepseek_V3_R1_...

Gemini CLI (blog.google)

U.S. bombs Iranian nuclear sites (bbc.co.uk)

Mechanical Watch: Exploded View (fellerts.no)

YouTube's new anti-adblock measures (iter.ca)

Samsung embeds IronSource spyware app on phones across WANA (smex.org)

Writing toy software is a joy (blog.jsbarretto.com)

uv: An extremely fast Python package and project manager, written in Rust (github.com)

OpenAI charges by the minute, so speed up your audio (george.mand.is)

Harper – an open-source alternative to Grammarly (writewithharper.com)

Phoenix.new – Remote AI Runtime for Phoenix (fly.io)

Fun with uv and PEP 723 (cottongeeks.com)

A new PNG spec (programmax.net)

Backyard Coffee and Jazz in Kyoto (thedeletedscenes.substack.com)

Git Notes: Git's coolest, most unloved­ feature (2022) (tylercipriani.com)

Vera C. Rubin Observatory first images (rubinobservatory.org)

A new pyramid-like shape always lands the same side up (quantamagazine.org)

A new PNG spec (programmax.net)

Man 'refused entry into US' as border control catch him with bald JD Vance meme (dublinlive.ie)

Thnickels (thick-coins.net)

How I use my terminal (jyn.dev)

I wrote my PhD Thesis in Typst (fransskarman.com)

Microsoft suspended the email account of an ICC prosecutor at The Hague (nytimes.com)

Define policy forbidding use of AI code generators (github.com)

Microsoft Edit (github.com)

-2000 Lines of code (folklore.org)

Hurl: Run and test HTTP requests with plain text (github.com)

Starship: A minimal, fast, and customizable prompt for any shell (starship.rs)

TPU Deep Dive (henryhmko.github.io)

PlasticList – Plastic Levels in Foods (plasticlist.org)

Klein Bottle Amazon Brand Hijacking (2021) (kleinbottle.com)

What Problems to Solve (1966) (genius.cat-v.org)

LibRedirect – Redirects popular sites to alternative privacy-friendly frontends (libredirect.github.io)

uBlock Origin Lite Beta for Safari iOS (testflight.apple.com)

Tell HN: Beware confidentiality agreements that act as lifetime non competes

Fairphone 6 is switching to a new design that's even more sustainable (androidcentral.com)

Finding a 27-year-old easter egg in the Power Mac G3 ROM (downtowndougbrown.com)

U.S. Chemical Safety Board could be eliminated (ishn.com)

Games run faster on SteamOS than Windows 11, Ars testing finds (arstechnica.com)

Using Home Assistant, adguard home and an $8 smart outlet to avoid brain rot (romanklasen.com)

Writing a basic Linux device driver when you know nothing about Linux drivers (crescentro.se)

GitHub CEO: manual coding remains key despite AI boom (techinasia.com)

New Linux udisks flaw lets attackers get root on major Linux distros (bleepingcomputer.com)

Basic Facts about GPUs (damek.github.io)

MCP is eating the world (stainless.com)

Python can run Mojo now (koaning.io)

AbsenceBench: Language models can't tell what's missing (arxiv.org)

Getting ready to issue IP address certificates (community.letsencrypt.org)

Show HN: Nxtscape – an open-source agentic browser (github.com)

Delta Chat is a decentralized and secure messenger app (delta.chat)

National Archives at College Park, MD, will become a restricted federal facility (archives.gov)

Ask HN: How Does DeepSeek "Thinks"?

Comments (2)

Git Notes: Git's coolest, most unloved feature (2022) (tylercipriani.com)