Describir: Comparative performance of large language models in structuring head CT radiology reports: multi-institutional validation study in Japan.