A Large-Scale Study of Programming Languages and Code Quality in GitHub.
What is the effect of programming languages on software quality? This question has been a topic of much debate for a very long time. In this study, we gather a very large data set from GitHub (728 projects, 63 million SLOC, 29,000 authors, 1.5 million commits, in 17 languages) in an attempt to shed...
| Publicado en: | Communications of the ACM Vol. 60; no. 10; pp. 91 - 101 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Association for Computing Machinery
Oct2017
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=125351984&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 125351984 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00010782 ACM jtl: Communications of the ACM issn: 00010782 maglogo: N pubinfo: dt: Oct2017 vid: 60 iid: 10 pid: 68 pub: Association for Computing Machinery artinfo: ui: 125351984 10.1145/3126905 ppf: 91 ppct: 10 formats: tig: atl: A Large-Scale Study of Programming Languages and Code Quality in GitHub. aug: au: Ray, Baishakhi Posnett, Daryl Posnett Devanbu, Premkumar Filkov, Vladimir affil: Department of Computer Science, University of Virginia, Charlottesville, VA. Department of Computer Science, University of California, Davis, CA. su: Programming languages Computer programming Computer software quality control Computer defects sug: subj: Programming languages Computer programming Computer software quality control Computer defects ab: What is the effect of programming languages on software quality? This question has been a topic of much debate for a very long time. In this study, we gather a very large data set from GitHub (728 projects, 63 million SLOC, 29,000 authors, 1.5 million commits, in 17 languages) in an attempt to shed some empirical light on this question. This reasonably large sample size allows us to use a mixed-methods approach, combining multiple regression modeling with visualization and text analytics, to study the effect of language features such as static versus dynamic typing and allowing versus disallowing type confusion on software quality. By triangulating findings from different methods, and controlling for confounding effects such as team size, project size, and project history, we report that language design does have a significant, but modest effect on software quality. Most notably, it does appear that disallowing type confusion is modestly better than allowing it, and among functional languages, static typing is also somewhat better than dynamic typing. We also find that functional languages are somewhat better than procedural languages. It is worth noting that these modest effects arising from language design are overwhelmingly dominated by the process factors such as project size, team size, and commit size. However, we caution the reader that even these modest effects might quite possibly be due to other, intangible process factors, for example, the preference of certain personality types for functional, static languages that disallow type confusion. pubtype: Periodical doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2017 holdings: @attributes: islocal: N |
|---|