Replymessage unavailable
That makes sense. For a service with confidential prompts and public model weights, what can a node operator actually see during inference: the prompt, derived embeddings or only encrypted transport metadata? That would make the design trade-off much clearer to me.