Describir: Blinded two-phase evaluation of large language models in complex cardiac surgery: task-specific performance and human-AI collaboration.