Ссылка
click to show
click to show
Two-Minute Papers just dropped a video on a mediation to one of the biggest safety concerns with AI.
Summary
Two-Minute Papers just dropped a video on a mediation to one of the biggest safety concerns with AI. Essentially, we are steering the AI back into the assistant role when it drifts off too hard, and could start being dangerous or unhelpful.
Quotes
Quote
"do it through an instant brain surgery" ... "Okay, got it. Now you take the brain activity when it is role-playing as a pirate, a goblin or something else. If you subtract the role player from the assistant, you get a vector. For simplicity, let's refer to this as helpfulness. Now let's keep our eye on helpfulness. If it goes below a threshold, we apply a nudge. How? Mathematically, we just measure how much helpfulness is in the model's thought. If it is above the safety line, fantastic. Keep watching as it works. But if it drops below the line, now that is trouble. The model is about to say
something inaccurate or dangerous. So now we calculate exactly how much is missing and add just enough helpfulness back into the equation. This pushes it back over the line. "
My thoughts
I've heard Luke and Linus talk about this many a time on WAN. I am glad to see they will be happy with the results coming along.
Also, note I did not say "fix"
Sources
Link