Browse Definitions :
Definition

Apache Lucene

Contributor(s): Stan Gibilisco

Apache Lucene is a freely available information retrieval software library that works with fields of text within document files. This evolving venture is also called the Apache Lucene Project. Apache is a server that is distributed under an open source license.

The Lucene application program interface (API) stays the same regardless of the format of the file to be indexed. Provided that the text information can be recovered and extracted, Lucene can index practically any type of text-containing document. Lucene has become popular for use in Internet search engines as well as for single-site search operations.

The Apache Lucene Project comprises four main components:

  • Lucene Core: Indexing, searching, spell checking, hit highlighting, and tokenization.
  • PyLucene: Python port for Lucene Core.
  • Solr: Extensible Markup Language (XML), Hypertext Transfer Protocol (HTTP), and APIs for Javascript Object Notation (JSON), Python, and Ruby, as well as hit highlighting, faceted search, caching, replication, and an interface for Web site administrators.
  • Open Relevance Project: Free distribution of materials for performance testing and relevance evaluation.
This was last updated in May 2013

Continue Reading About Apache Lucene

Join the conversation

1 comment

Send me notifications when other members comment.

Please create a username to comment.

can somebody please explain about what is the meaning of ranking in lucene. why we use it.
Cancel

-ADS BY GOOGLE

File Extensions and File Formats

SearchCompliance

  • PCI DSS (Payment Card Industry Data Security Standard)

    The Payment Card Industry Data Security Standard (PCI DSS) is a widely accepted set of policies and procedures intended to ...

  • risk management

    Risk management is the process of identifying, assessing and controlling threats to an organization's capital and earnings.

  • compliance framework

    A compliance framework is a structured set of guidelines that details an organization's processes for maintaining accordance with...

SearchSecurity

  • Trojan horse (computing)

    In computing, a Trojan horse is a program downloaded and installed on a computer that appears harmless, but is, in fact, ...

  • identity theft

    Identity theft, also known as identity fraud, is a crime in which an imposter obtains key pieces of personally identifiable ...

  • DNS over HTTPS (DoH)

    DNS over HTTPS (DoH) is a relatively new protocol that encrypts domain name system traffic by passing DNS queries through a ...

SearchHealthIT

  • telemedicine (telehealth)

    Telemedicine is the remote delivery of healthcare services, such as health assessments or consultations, over the ...

  • Project Nightingale

    Project Nightingale is a controversial partnership between Google and Ascension, the second largest health system in the United ...

  • medical practice management (MPM) software

    Medical practice management (MPM) software is a collection of computerized services used by healthcare professionals and ...

SearchDisasterRecovery

SearchStorage

  • M.2 SSD

    An M.2 SSD is a solid-state drive (SSD) that conforms to a computer industry specification and is used in internally mounted ...

  • kilobyte (KB or Kbyte)

    A kilobyte (KB or Kbyte) is a unit of measurement for computer memory or data storage used by mathematics and computer science ...

  • virtual memory

    Virtual memory is a memory management capability of an operating system (OS) that uses hardware and software to allow a computer ...

Close