江恩买卖十二法则word:Lucene排序算法

来源：百度文库编辑：中财网时间：2024/04/28 08:44:41

The score of query q for document d correlates to thecosine-distance or dot-product between document and query vectors in aVector Space Model (VSM) of Information Retrieval.A document whose vector is closer to the query vector in that model is scored higher.The score is computed as follows:

score(q,d) = coord(q,d) · queryNorm(q) · ∑ ( tf(t in d) · idf(t)² · t.getBoost() · norm(t,d) )
t in q

where

tf(t in d) correlates to the term's frequency, defined as the number of times term t appears in the currently scored document d. Documents that have more occurrences of a given term receive a higher score. The default computation for tf(t in d) in DefaultSimilarity is:

tf(t in d) = frequency^½
idf(t) stands for Inverse Document Frequency. This value correlates to the inverse of docFreq (the number of documents in which the term t appears). This means rarer terms give higher contribution to the total score. The default computation for idf(t) in DefaultSimilarity is:

idf(t) = 1 + log ( numDocs ––––––––– docFreq+1 )
coord(q,d) is a score factor based on how many of the query terms are found in the specified document. Typically, a document that contains more of the query's terms will receive a higher score than another document with fewer query terms. This is a search time factor computed in coord(q,d) by the Similarity in effect at search time.
queryNorm(q) is a normalizing factor used to make scores between queries comparable. This factor does not affect document ranking (since all ranked documents are multiplied by the same factor), but rather just attempts to make scores from different queries (or even different indexes) comparable. This is a search time factor computed by the Similarity in effect at search time. The default computation in DefaultSimilarity is:

queryNorm(q) = queryNorm(sumOfSquaredWeights) = 1 –––––––––––––– sumOfSquaredWeights^½

The sum of squared weights (of the query terms) is computed by the query Weight object. For example, a boolean query computes this value as:

sumOfSquaredWeights = q.getBoost() ² · ∑ ( idf(t) · t.getBoost() ) ²
t in q
t.getBoost() is a search time boost of term t in the query q as specified in the query text (see query syntax), or as set by application calls to setBoost(). Notice that there is really no direct API for accessing a boost of one term in a multi term query, but rather multi terms are represented in a query as multi TermQuery objects, and so the boost of a term in the query is accessible by calling the sub-query getBoost().
norm(t,d) encapsulates a few (indexing time) boost and length factors:
- Document boost - set by calling doc.setBoost() before adding the document to the index.
- Field boost - set by calling field.setBoost() before adding the field to a document.
- lengthNorm(field) - computed when the document is added to the index in accordance with the number of tokens of this field in the document, so that shorter fields contribute more to the score. LengthNorm is computed by the Similarity class in effect at indexing.
When a document is added to the index, all the above factors are multiplied. If the document has multiple fields with the same name, all their boosts are multiplied together:

norm(t,d) = doc.getBoost() · lengthNorm(field) · ∏ f.getBoost()
field f in d named as t

However the resulted norm value is encoded as a single byte before being stored. At search time, the norm byte value is read from the index directory and decoded back to a float norm value. This encoding/decoding, while reducing index size, comes with the price of precision loss - it is not guaranteed that decode(encode(x)) = x. For instance, decode(encode(0.89)) = 0.75. Also notice that search time is too late to modify this norm part of scoring, e.g. by using a different Similarity for search.

结构体排序算法 asp排序算法什么是起泡排序算法?? 急求一排序算法合并排序算法 vb选择排序算法内部排序算法比较内部排序算法比较整数排序算法的问题？堆排序的具体算法快速排序的循环算法求几种排序法的算法全排序的算法(PASCAL) 请教算法高手：还是关于排序！ pascal 排序算法的一道选择题。Help！给出一种拓扑排序的改进算法关于数据库排序输出的算法? 用C语言写一个排序算法。 C语言中排序的算法是什么用JAVA写出冒泡排序的算法链表的冒泡排序算法关于排序和查找的算法用pascal语言实现希尔排序算法各种排序算法实现和比较