SearcharxivSearch

arXiv subjects

Dirk Lewandowski

Publications and source records attributed to Dirk Lewandowski.

At least 19 recordsLinked to original sources

Beyond Topicality: A Conceptual Analysis of Societal Relevance and Its Application to Search Results and AI Responses

This paper examines "societal relevance," a concept introduced by Haider and Sundin to address the limitations of traditional relevance models in web search. While topical and user relevance are foundational to information science, they are insufficient for managing harmful content such as misinformation or discrimination found on the uncontrolled web. This study investigates three analytical questions: the definition of societal relevance, its practical application in search systems, and its distinction from information quality measures. By analyzing various combinations of system, user, and societal relevance, the paper explores how search outputs can be optimized for the "greater good". Although the concept remains theoretically underdeveloped, it provides a vital framework for developing value-driven search engines that prioritize ethical outcomes and societal interests over mere keyword matching.

cs.IR

Is googling risky? A study on risk perception and experiences of adverse consequences in web search

Search engines, such as Google, have a considerable impact on society. Therefore, undesirable consequences, such as retrieving incorrect search results, pose a risk to users. Although previous research has reported the adverse outcomes of web search, little is known about how search engine users evaluate those outcomes. In this study, we show which aspects of web search are perceived as risky using a sample (N = 3,884) representative of the German Internet population. We found that many participants are often concerned with adverse consequences immediately appearing on the search engine result page. Moreover, participants' experiences with adverse consequences are directly related to their risk perception. Our results demonstrate that people perceive risks related to web search. In addition to our study, there is a need for more independent research on the possible detrimental outcomes of web search to monitor and mitigate risks. Apart from risks for individuals, search engines with a massive number of users have an extraordinary impact on society; therefore, the acceptable risks of web search should be discussed.

cs.IR

Does Search Engine Optimization come along with high-quality content? A comparison between optimized and non-optimized health-related web pages

Searching for medical information is both a common and important activity since it influences decisions people make about their healthcare. Using search engine optimization (SEO), content producers seek to increase the visibility of their content. SEO is more likely to be practiced by commercially motivated content producers such as pharmaceutical companies than by non-commercial providers such as governmental bodies. In this study, we ask whether content quality correlates with the presence or absence of SEO measures on a web page. We conducted a user study in which N = 61 participants comprising laypeople as well as experts in health information assessment evaluated health-related web pages classified as either optimized or non-optimized. The subjects rated the expertise of non-optimized web pages as higher than the expertise of optimized pages, justifying their appraisal by the more competent and reputable appearance of non-optimized pages. In addition, comments about the website operators of the non-optimized pages were exclusively positive, while optimized pages tended to receive positive as well as negative assessments. We found no differences between the ratings of laypeople and experts. Since non-optimized, but high-quality content may be outranked by optimized content of lower quality, trusted sources should be prioritized in rankings.

cs.IR

A Comparison of Source Distribution and Result Overlap in Web Search Engines

When it comes to search engines, users generally prefer Google. Our study aims to find the differences between the results found in Google compared to other search engines. We compared the top 10 results from Google, Bing, DuckDuckGo, and Metager, using 3,537 queries generated from Google Trends from Germany and the US. Google displays more unique domains in the top results than its competitors. Wikipedia and news websites are the most popular sources overall. With some top sources dominating search results, the distribution of domains is also consistent across all search engines. The overlap between Google and Bing is always under 32%, while Metager has a higher overlap with Bing than DuckDuckGo, going up to 78%. This study shows that the use of another search engine, especially in addition to Google, provides a wider variety in sources and might lead the user to find new perspectives.

cs.IR

Public awareness and attitudes towards search engine optimization

This research focuses on what users know about search engine optimization (SEO) and how well they can identify results that have potentially been influenced by SEO. We conducted an online survey with a sample representative of the German online population (N = 2,012). We found that 43% of users assume a better ranking can be achieved without paying money to Google. This is in stark contrast to the possibility of influence through paid advertisements, which 79% of internet users are aware of. However, only 29.2% know how ads differ from organic results. The term "search engine optimization" is known to 8.9% of users but 14.5% can correctly name at least one SEO tactic. Success in labelling results that can be influenced through SEO varies by search engine result page (SERP) complexity and devices: participants achieved higher success rates on SERPs with simple structures than on the more complex SERPs. SEO results were identified better on the small screen than on the large screen. 59.2% assumed that SEO has a (very) strong impact on rankings. SEO is more often perceived as positive (75.2%) than negative (68.4%). The insights from this study have implications for search engine providers, regulators, and information literacy.

cs.IR

Misplaced trust? The relationship between trust, ability to identify commercially influenced results, and search engine preference

People have a high level of trust in search engines, especially Google, but only limited knowledge of them, as numerous studies have shown. This leads to the question: To what extent is this trust justified considering the lack of familiarity among users with how search engines work and the business models they are founded on? We assume that trust in Google, search engine preferences, and knowledge of result types are interrelated. To examine this assumption, we conducted a representative online survey with n = 2,012 German internet users. We show that users with little search engine knowledge are more likely to trust and use Google than users with more knowledge. A contradiction revealed itself - users strongly trust Google, yet they are unable to adequately evaluate search results. This may be problematic since it can potentially affect knowledge acquisition. Consequently, there is a need to promote user information literacy to create a more solid foundation for user trust in search engines.

cs.IR

An Empirical Investigation On Search Engine Ad Disclosure

This representative study of German search engine users (N=1,000) focuses on the ability of users to distinguish between organic results and advertisements on Google results pages. We combine questions about Google's business with task-based studies in which users were asked to distinguish between ads and organic results in screenshots of results pages. We find that only a small percentage of users is able to reliably distinguish between ads and organic results, and that user knowledge of Google's business model is very limited. We conclude that ads are insufficiently labelled as such, and that many users may click on ads assuming that they are selecting organic results.

cs.IR

Does it matter which search engine is used? A user study using post-task relevance judgments

The objective of this research was to find out how the two search engines Google and Bing perform when users work freely on pre-defined tasks, and judge the relevance of the results immediately after finishing their search session. In a user study, 64 participants conducted two search tasks each, and then judged the results on the following: (1) The quality of the results they selected in their search sessions, (2) The quality of the results they were presented with in their search sessions (but which they did not click on), (3) The quality of the results from the competing search engine for their queries (which they did not see in their search session). We found that users heavily relied on Google, that Google produced more relevant results than Bing, that users were well able to select relevant results from the results lists, and that users judged the relevance of results lower when they regarded a task as difficult and did not find the correct information.

cs.HC

How Relevant is the Long Tail? A Relevance Assessment Study on Million Short

Users of web search engines are known to mostly focus on the top ranked results of the search engine result page. While many studies support this well known information seeking pattern only few studies concentrate on the question what users are missing by neglecting lower ranked results. To learn more about the relevance distributions in the so-called long tail we conducted a relevance assessment study with the Million Short long-tail web search engine. While we see a clear difference in the content between the head and the tail of the search engine result list we see no statistical significant differences in the binary relevance judgments and weak significant differences when using graded relevance. The tail contains different but still valuable results. We argue that the long tail can be a rich source for the diversification of web search engine result lists but it needs more evaluation to clearly describe the differences.

cs.IR

Problems with the use of Web search engines to find results in foreign languages

Purpose - To test the ability of major search engines, Google, Yahoo, MSN, and Ask, to distinguish between German and English-language documents Design/methodology/approach - 50 queries, using words common in German and in English, were posed to the engines. The advanced search option of language restriction was used, once in German and once in English. The first 20 results per engine in each language were investigated. Findings - While none of the search engines faces problems in providing results in the language of the interface that is used, both Google and MSN face problems when the results are restricted to a foreign language. Research limitations/implications - Search engines were only tested in German and in English. We have only anecdotal evidence that the problems are the same with other languages. Practical implications - Searchers should not use the language restriction in Google and MSN when searching for foreign-language documents. Instead, searchers should use Yahoo or Ask. If searching for foreign language documents in Google or MSN, the interface in the target language/country should be used. Value of paper - Demonstrates a problem with search engines that has not been previously investigated.

cs.IR

The Retrieval Effectiveness of Web Search Engines: Considering Results Descriptions

Purpose: To compare five major Web search engines (Google, Yahoo, MSN, Ask.com, and Seekport) for their retrieval effectiveness, taking into account not only the results but also the results descriptions. Design/Methodology/Approach: The study uses real-life queries. Results are made anonymous and are randomised. Results are judged by the persons posing the original queries. Findings: The two major search engines, Google and Yahoo, perform best, and there are no significant differences between them. Google delivers significantly more relevant result descriptions than any other search engine. This could be one reason for users perceiving this engine as superior. Research Limitations: The study is based on a user model where the user takes into account a certain amount of results rather systematically. This may not be the case in real life. Practical Implications: Implies that search engines should focus on relevant descriptions. Searchers are advised to use other search engines in addition to Google. Originality/Value: This is the first major study comparing results and descriptions systematically and proposes new retrieval measures to take into account results descriptions

cs.IR

What Users See - Structures in Search Engine Results Pages

This paper investigates the composition of search engine results pages. We define what elements the most popular web search engines use on their results pages (e.g., organic results, advertisements, shortcuts) and to which degree they are used for popular vs. rare queries. Therefore, we send 500 queries of both types to the major search engines Google, Yahoo, Live.com and Ask. We count how often the different elements are used by the individual engines. In total, our study is based on 42,758 elements. Findings include that search engines use quite different approaches to results pages composition and therefore, the user gets to see quite different results sets depending on the search engine and search query used. Organic results still play the major role in the results pages, but different shortcuts are of some importance, too. Regarding the frequency of certain host within the results sets, we find that all search engines show Wikipedia results quite often, while other hosts shown depend on the search engine used. Both Google and Yahoo prefer results from their own offerings (such as YouTube or Yahoo Answers). Since we used the .com interfaces of the search engines, results may not be valid for other country-specific interfaces.

cs.IR

Ranking library materials

Purpose: This paper discusses ranking factors suitable for library materials and shows that ranking in general is a complex process and that ranking for library materials requires a variety of techniques. Design/methodology/approach: The relevant literature is reviewed to provide a systematic overview of suitable ranking factors. The discussion is based on an overview of ranking factors used in Web search engines. Findings: While there are a wide variety of ranking factors applicable to library materials, todays library systems use only some of them. When designing a ranking component for the library catalogue, an individual weighting of applicable factors is necessary. Research limitations/applications: While this article discusses different factors, no particular ranking formula is given. However, this article presents the argument that such a formula must always be individual to a certain use case. Practical implications: The factors presented can be considered when designing a ranking component for a librarys search system or when discussing such a project with an ILS vendor. Originality/value: This paper is original in that it is the first to systematically discuss ranking of library materials based on the main factors used by Web search engines.

cs.DL

Using Search Engine Technology to Improve Library Catalogs

This chapter outlines how search engine technology can be used in online public access library catalogs (OPACs) to help improve users experiences, to identify users intentions, and to indicate how it can be applied in the library context, along with how sophisticated ranking criteria can be applied to the online library catalog. A review of the literature and current OPAC developments form the basis of recommendations on how to improve OPACs. Findings were that the major shortcomings of current OPACs are that they are not sufficiently user-centered and that their results presentations lack sophistication. Further, these shortcomings are not addressed in current 2.0 developments. It is argued that OPAC development should be made search-centered before additional features are applied. While the recommendations on ranking functionality and the use of user intentions are only conceptual and not yet applied to a library catalogue, practitioners will find recommendations for developing better OPACs in this chapter. In short, readers will find a systematic view on how the search engines strengths can be applied to improving libraries online catalogs.

cs.DL

Google Scholar as a tool for discovering journal articles in library and information science

Purpose: The purpose of this paper is to measure the coverage of Google Scholar for the Library and Information Science (LIS) journal literature as identified by a list of core LIS journals from a study by Schloegl and Petschnig (2005). Methods: We checked every article from 35 major LIS journals from the years 2004 to 2006 for availability in Google Scholar (GS). We also collected information on the type of availability-i.e., whether a certain article was available as a PDF for a fee, as a free PDF, or as a preprint. Results: We found that only some journals are completely indexed by Google Scholar, that the ratio of versions available depends on the type of publisher, and that availability varies a lot from journal to journal. Google Scholar cannot substitute for abstracting and indexing services in that it does not cover the complete literature of the field. However, it can be used in many cases to easily find available full texts of articles already found using another tool. Originality/value: This study differs from other Google Scholar coverage studies in that it takes into account not only whether an article is indexed in GS at all, but also the type of availability.

cs.DL

The Influence of Commercial Intent of Search Results on Their Perceived Relevance

We carried out a retrieval effectiveness test on the three major web search engines (i.e., Google, Microsoft and Yahoo). In addition to relevance judgments, we classified the results according to their commercial intent and whether or not they carried any advertising. We found that all search engines provide a large number of results with a commercial intent. Google provides significantly more commercial results than the other search engines do. However, the commercial intent of a result did not influence jurors in their relevance judgments.

cs.IR

The retrieval effectiveness of search engines on navigational queries

Purpose - To test major Web search engines on their performance on navigational queries, i.e. searches for homepages. Design/methodology/approach - 100 real user queries are posed to six search engines (Google, Yahoo, MSN, Ask, Seekport, and Exalead). Users described the desired pages, and the results position of these is recorded. Measured success N and mean reciprocal rank are calculated. Findings - Performance of the major search engines Google, Yahoo, and MSN is best, with around 90 percent of queries answered correctly. Ask and Exalead perform worse but receive good scores as well. Research limitations/implications - All queries were in German, and the German-language interfaces of the search engines were used. Therefore, the results are only valid for German queries. Practical implications - When designing a search engine to compete with the major search engines, care should be taken on the performance on navigational queries. Users can be influenced easily in their quality ratings of search engines based on this performance. Originality/value - This study systematically compares the major search engines on navigational queries and compares the findings with studies on the retrieval effectiveness of the engines on informational queries. Paper type - research paper

cs.IR