ChatCrawlersearch across public Telegram Open the app
L

Longevity InTime: Autonomous AI Institute. Anti-Aging Digital Health Immortality Transhumanist AI Channel

сообщение · 2026-09-02 20:08 UTC
C
Ссылка
click to show
Grok 4.6 Scores Over 50% on Both Safety and Utility Scales in BioSecBench-Refusal Test LatchBio, the developer of the BioSecBench-Refusal test, reported on 1 September that its evaluation of Grok 4.6 showed the model exceeding 50% on both safety and utility scales. The company said Grok 4.6 is the only model tested so far to achieve this result independently of the software wrapper used to run it. The BioSecBench-Refusal benchmark evaluates an AI agent on two kinds of biological tasks. In 61 ordinary tasks the agent must assist with data analysis from published studies, while in 46 dangerous scenarios a harmful intent is concealed in the input or the task description. Because the same biological terminology appears in both sets, the test records two error types: missing a hidden threat and refusing legitimate research work. In the July check of BioSecBench-Refusal most of the 16 model‑wrapper combinations refused ordinary tasks at least as often as they refused risky scenarios, with filters often cutting off work before the agent could parse the goal. The new series measures both sides of this problem specifically for Grok 4.6. LatchBio attributes the score to Grok 4.6’s ability to match a harmless cover story with the content of an attached file and to consider the experiment’s purpose. In their examples a request to neutralize a 🔗 Read original →
15 ·

Вся лента · оригинал в Telegram

Open in Telegram Feed t.me/s Каталог площадок Искать в ChatCrawler

A snapshot of an open public feed from the search index ChatCrawler — “Google for public Telegram”; refreshed as the venue is crawled. Times are UTC.

Public content only, official Telegram API. About · FAQ · What we do not do · Remove a page · Catalog