r/LocalLLaMA

11 stories

r/LocalLLaMA is a Reddit community focused on running large language models locally and related tooling. Recent coverage highlights Agnes-3.0-Flash (33B multimodal model with a 262,144-token context window and hybrid-attention decoder), flash models reaching 512GB, a user switching to open weights after OpenAI blocked a protein-design project, Deepseek soft-retiring V4 Pro, and community guidance as the subreddit grows.

Related topics

Agnes-3.0-Flash multimodal model

Agnes-3.0-Flash, posted on r/LocalLLaMA, is a 33B multimodal model with a 262,144-token context window. It uses a hybrid-attention decoder (72 layers: 54 delta-rule recurrent + 18 global-attention), includes vision/video understanding, and targets long-context and tool-calling use cases.

r/LocalLLaMA · · Details

Notes on a Hobby Sub Going Mainstream

A longtime user outlines pros and cons of a hobby subreddit growing mainstream, advising newcomers to stay humble, keep politics out unless policy-related, and avoid repeating exhausted debates to preserve a technical, evidence-based community.

r/LocalLLaMA · · Details
That is everything