Re: [PATCH 2/5] util: Function to parse boolean term queries
authorAustin Clements <amdragon@MIT.EDU>
Wed, 26 Dec 2012 01:23:00 +0000 (20:23 +1900)
committerW. Trevor King <wking@tremily.us>
Fri, 7 Nov 2014 17:52:48 +0000 (09:52 -0800)
14/f4d0faaedb51c7afd590d7ebf4267a6d59c17b [new file with mode: 0644]

diff --git a/14/f4d0faaedb51c7afd590d7ebf4267a6d59c17b b/14/f4d0faaedb51c7afd590d7ebf4267a6d59c17b
new file mode 100644 (file)
index 0000000..8f16822
--- /dev/null
@@ -0,0 +1,152 @@
+Return-Path: <amdragon@mit.edu>\r
+X-Original-To: notmuch@notmuchmail.org\r
+Delivered-To: notmuch@notmuchmail.org\r
+Received: from localhost (localhost [127.0.0.1])\r
+       by olra.theworths.org (Postfix) with ESMTP id 9D5AD431FAF\r
+       for <notmuch@notmuchmail.org>; Tue, 25 Dec 2012 17:23:07 -0800 (PST)\r
+X-Virus-Scanned: Debian amavisd-new at olra.theworths.org\r
+X-Spam-Flag: NO\r
+X-Spam-Score: -0.7\r
+X-Spam-Level: \r
+X-Spam-Status: No, score=-0.7 tagged_above=-999 required=5\r
+       tests=[RCVD_IN_DNSWL_LOW=-0.7] autolearn=disabled\r
+Received: from olra.theworths.org ([127.0.0.1])\r
+       by localhost (olra.theworths.org [127.0.0.1]) (amavisd-new, port 10024)\r
+       with ESMTP id m+ueMU2URmHt for <notmuch@notmuchmail.org>;\r
+       Tue, 25 Dec 2012 17:23:07 -0800 (PST)\r
+Received: from dmz-mailsec-scanner-3.mit.edu (DMZ-MAILSEC-SCANNER-3.MIT.EDU\r
+       [18.9.25.14])\r
+       by olra.theworths.org (Postfix) with ESMTP id DC466431FAE\r
+       for <notmuch@notmuchmail.org>; Tue, 25 Dec 2012 17:23:06 -0800 (PST)\r
+X-AuditID: 1209190e-b7fa16d000001402-81-50da51791370\r
+Received: from mailhub-auth-3.mit.edu ( [18.9.21.43])\r
+       by dmz-mailsec-scanner-3.mit.edu (Symantec Messaging Gateway) with SMTP\r
+       id 5E.0D.05122.9715AD05; Tue, 25 Dec 2012 20:23:05 -0500 (EST)\r
+Received: from outgoing.mit.edu (OUTGOING-AUTH.MIT.EDU [18.7.22.103])\r
+       by mailhub-auth-3.mit.edu (8.13.8/8.9.2) with ESMTP id qBQ1N43n016953; \r
+       Tue, 25 Dec 2012 20:23:04 -0500\r
+Received: from awakening.csail.mit.edu (awakening.csail.mit.edu [18.26.4.91])\r
+       (authenticated bits=0)\r
+       (User authenticated as amdragon@ATHENA.MIT.EDU)\r
+       by outgoing.mit.edu (8.13.6/8.12.4) with ESMTP id qBQ1N1HD002404\r
+       (version=TLSv1/SSLv3 cipher=DHE-RSA-AES128-SHA bits=128 verify=NOT);\r
+       Tue, 25 Dec 2012 20:23:02 -0500 (EST)\r
+Received: from amthrax by awakening.csail.mit.edu with local (Exim 4.80)\r
+       (envelope-from <amdragon@MIT.EDU>)\r
+       id 1Tnfi0-0008Ut-Qk; Tue, 25 Dec 2012 20:23:00 -0500\r
+Date: Tue, 25 Dec 2012 20:23:00 -0500\r
+From: Austin Clements <amdragon@MIT.EDU>\r
+To: David Bremner <david@tethera.net>\r
+Subject: Re: [PATCH 2/5] util: Function to parse boolean term queries\r
+Message-ID: <20121226012300.GW6187@mit.edu>\r
+References: <1356415076-5692-1-git-send-email-amdragon@mit.edu>\r
+       <1356415076-5692-3-git-send-email-amdragon@mit.edu>\r
+       <87obhidxkt.fsf@zancas.localnet>\r
+MIME-Version: 1.0\r
+Content-Type: text/plain; charset=us-ascii\r
+Content-Disposition: inline\r
+In-Reply-To: <87obhidxkt.fsf@zancas.localnet>\r
+User-Agent: Mutt/1.5.21 (2010-09-15)\r
+X-Brightmail-Tracker:\r
+ H4sIAAAAAAAAA+NgFprOKsWRmVeSWpSXmKPExsUixCmqrVsZeCvAYMNMFYsbrd2MFk3TnS1W\r
+       z+WxuH5zJrMDi8fOWXfZPW7df83u8WzVLWaPLYfeMwewRHHZpKTmZJalFunbJXBldHxbwFbQ\r
+       KVzx/vZa1gbGDr4uRk4OCQETibZfTUwQtpjEhXvr2boYuTiEBPYxSnzpnMgO4WxglLi1qAPK\r
+       ucgk0bPpOzOEs4RRYtKdh2D9LAKqEjdnvWcEsdkENCS27V8OZosAxa9um8wGYjMLuEusn3gG\r
+       rF5YwFXifmMvWA2vgLbE5+tTGSGGzmCUePRxFTtEQlDi5MwnLBDNWhI3/r0EauYAsqUllv/j\r
+       AAlzCuhKNP4+BzZTVEBFYsrJbWwTGIVmIemehaR7FkL3AkbmVYyyKblVurmJmTnFqcm6xcmJ\r
+       eXmpRbrGermZJXqpKaWbGMHhL8m3g/HrQaVDjAIcjEo8vBu/3wwQYk0sK67MPcQoycGkJMp7\r
+       3v9WgBBfUn5KZUZicUZ8UWlOavEhRgkOZiURXuePQOW8KYmVValF+TApaQ4WJXHeKyk3/YUE\r
+       0hNLUrNTUwtSi2CyMhwcShK8cwOAhgoWpaanVqRl5pQgpJk4OEGG8wANnwpSw1tckJhbnJkO\r
+       kT/FqCglzpsLkhAASWSU5sH1wtLTK0ZxoFeEeZeBVPEAUxtc9yugwUxAg2P5boAMLklESEk1\r
+       ME73DVxwvsr5by/vJsFPWXd/yM+ePuHeuYKlrpPnqZTZSd9dZKuRz5JzMe/5JtEQv+X8eZzb\r
+       lCYb6+4yrtjJpp9wrpXhjWzMxbK38VtX3zrCkvtisl3prO8TXqfpBN3eKlpt6lFXOpNr6tag\r
+       d/H+a39N/N9wfQHLUXOjC2d9a+3cbQ6pWM0WdVNiKc5INNRiLipOBAAJELYQKgMAAA==\r
+Cc: notmuch@notmuchmail.org\r
+X-BeenThere: notmuch@notmuchmail.org\r
+X-Mailman-Version: 2.1.13\r
+Precedence: list\r
+List-Id: "Use and development of the notmuch mail system."\r
+       <notmuch.notmuchmail.org>\r
+List-Unsubscribe: <http://notmuchmail.org/mailman/options/notmuch>,\r
+       <mailto:notmuch-request@notmuchmail.org?subject=unsubscribe>\r
+List-Archive: <http://notmuchmail.org/pipermail/notmuch>\r
+List-Post: <mailto:notmuch@notmuchmail.org>\r
+List-Help: <mailto:notmuch-request@notmuchmail.org?subject=help>\r
+List-Subscribe: <http://notmuchmail.org/mailman/listinfo/notmuch>,\r
+       <mailto:notmuch-request@notmuchmail.org?subject=subscribe>\r
+X-List-Received-Date: Wed, 26 Dec 2012 01:23:07 -0000\r
+\r
+Quoth David Bremner on Dec 25 at 10:18 am:\r
+> Austin Clements <amdragon@MIT.EDU> writes:\r
+> \r
+> > +    if (consume_double_quote (&pos)) {\r
+> > +  char *out = talloc_strdup (ctx, pos);\r
+> > +  pos = *term_out = out;\r
+> > +  while (1) {\r
+> \r
+> Overall the control flow here is a bit tricky to follow. I'm not sure if\r
+> a real loop condition would help or make it worse.\r
+> \r
+> > +      if (! *pos) {\r
+> > +          /* Premature end of string */\r
+> > +          goto FAIL;\r
+> > +      } else if (*pos == '"') {\r
+> > +          if (*++pos != '"')\r
+> > +              break;\r
+> > +      } else if (consume_double_quote (&pos)) {\r
+> > +          break;\r
+> > +      }\r
+> \r
+> I'm confused by the asymmetry here. Quoted strings can start with\r
+> unicode quotes, but internally can only have ascii '"'? Is this\r
+> documented somewhere?\r
+\r
+Only in the source, to my knowledge.  Here's how Xapian parses these\r
+things (where 'it' is a UTF8 string iterator):\r
+\r
+if (it != end && is_double_quote(*it)) {\r
+    // Quoted boolean term (can contain any character).\r
+    ++it;\r
+    while (it != end) {\r
+       if (*it == '"') {\r
+           // Interpret "" as an escaped ".\r
+           if (++it == end || *it != '"')\r
+               break;\r
+       } else if (is_double_quote(*it)) {\r
+           ++it;\r
+           break;\r
+       }\r
+       Unicode::append_utf8(name, *it++);\r
+    }\r
+} else {\r
+    // Can't boolean filter prefix a subexpression, so\r
+    // just use anything following the prefix until the\r
+    // next space or ')' as part of the boolean filter\r
+    // term.\r
+    while (it != end && *it > ' ' && *it != ')')\r
+       Unicode::append_utf8(name, *it++);\r
+}\r
+\r
+> > +    } else {\r
+> > +  while (*pos > ' ' && *pos != ')')\r
+> > +      ++pos;\r
+> > +  if (*pos)\r
+> > +      goto FAIL;\r
+> > +    }\r
+> \r
+> So if there is no quote, we skip the part after the ':'? I guess I\r
+> probably missed something because that doesn't sound like the intended\r
+> behaviour.\r
+\r
+This isn't skipping it; it's checking its well-formedness.  In this\r
+case, *term_out already points to a correct string that can be used\r
+literally; we just have to check that there's no trailing garbage\r
+after the boolean query.\r
+\r
+This is certainly worth commenting.\r
+\r
+For the record, I also tried passing the query straight to the\r
+library, without parsing it in the CLI (and simply checking that the\r
+query returned exactly one result), and it was noticeably slower (the\r
+restore performance test took between 3.82 and 5.25 seconds for the\r
+code in this series and ~7.2 seconds using a general query.)\r