Can we use Google Scholar to identify highly-cited documents?

Alberto Martin-martin, Enrique Orduna-malea, Anne-wil Harzing, Emilio Delgado López-cózar

Research output: Contribution to journalArticleScientificpeer-review

104 Citations (Scopus)


The main objective of this paper is to empirically test whether the identification of highly-cited documents through Google Scholar is feasible and reliable. To this end, we carried out a longitudinal analysis (1950–2013), running a generic query (filtered only by year of publication) to minimise the effects of academic search engine optimisation. This gave us a final sample of 64,000 documents (1000 per year). The strong correlation between a document’s citations and its position in the search results (r = −0.67) led us to conclude that Google Scholar is able to identify highly-cited papers effectively. This, combined with Google Scholar’s unique coverage (no restrictions on document type and source), makes the academic search engine an invaluable tool for bibliometric research relating to the identification of the most influential scientific documents. We find evidence, however, that Google Scholar ranks those documents whose language (or geographical web domain) matches with the user’s interface language higher than could be expected based on citations. Nonetheless, this language effect and other factors related to the Google Scholar’s operation, i.e. the proper identification of versions and the date of publication, only have an incidental impact. They do not compromise the ability of Google Scholar to identify the highly-cited papers.
Original languageEnglish
Pages (from-to)152-163
JournalJournal of Informetrics
Issue number1
Publication statusPublished - 1 Feb 2017
Externally publishedYes


  • Google Scholar
  • academic search engines
  • highly-cited documents
  • academic information retrieval


Dive into the research topics of 'Can we use Google Scholar to identify highly-cited documents?'. Together they form a unique fingerprint.

Cite this