Re: [RFC PATCH 00/14] modular mail stores based on URIs
authorEthan <ethan.glasser.camp@gmail.com>
Sun, 1 Jul 2012 16:02:08 +0000 (12:02 +2000)
committerW. Trevor King <wking@tremily.us>
Fri, 7 Nov 2014 17:47:54 +0000 (09:47 -0800)
44/eeeeb104b373d62be3a90e88430dfe8d96606f [new file with mode: 0644]

diff --git a/44/eeeeb104b373d62be3a90e88430dfe8d96606f b/44/eeeeb104b373d62be3a90e88430dfe8d96606f
new file mode 100644 (file)
index 0000000..5ff3aa2
--- /dev/null
@@ -0,0 +1,200 @@
+Return-Path: <ethan.glasser.camp@gmail.com>\r
+X-Original-To: notmuch@notmuchmail.org\r
+Delivered-To: notmuch@notmuchmail.org\r
+Received: from localhost (localhost [127.0.0.1])\r
+       by olra.theworths.org (Postfix) with ESMTP id 49766431FAF\r
+       for <notmuch@notmuchmail.org>; Sun,  1 Jul 2012 09:02:12 -0700 (PDT)\r
+X-Virus-Scanned: Debian amavisd-new at olra.theworths.org\r
+X-Spam-Flag: NO\r
+X-Spam-Score: -0.798\r
+X-Spam-Level: \r
+X-Spam-Status: No, score=-0.798 tagged_above=-999 required=5\r
+       tests=[DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1,\r
+       FREEMAIL_FROM=0.001, HTML_MESSAGE=0.001, RCVD_IN_DNSWL_LOW=-0.7]\r
+       autolearn=disabled\r
+Received: from olra.theworths.org ([127.0.0.1])\r
+       by localhost (olra.theworths.org [127.0.0.1]) (amavisd-new, port 10024)\r
+       with ESMTP id ppDft23gV5PP for <notmuch@notmuchmail.org>;\r
+       Sun,  1 Jul 2012 09:02:10 -0700 (PDT)\r
+Received: from mail-vc0-f181.google.com (mail-vc0-f181.google.com\r
+       [209.85.220.181]) (using TLSv1 with cipher RC4-SHA (128/128 bits))\r
+       (No client certificate requested)\r
+       by olra.theworths.org (Postfix) with ESMTPS id EE8B7431FAE\r
+       for <notmuch@notmuchmail.org>; Sun,  1 Jul 2012 09:02:09 -0700 (PDT)\r
+Received: by vcbf1 with SMTP id f1so3668325vcb.26\r
+       for <notmuch@notmuchmail.org>; Sun, 01 Jul 2012 09:02:08 -0700 (PDT)\r
+DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113;\r
+       h=mime-version:in-reply-to:references:date:message-id:subject:from:to\r
+       :cc:content-type;\r
+       bh=UeoCrT7iWqZ6Z/rdq9M9nP8td9UW6zCNti7yveOS2lo=;\r
+       b=TX0jkZbJQgm7exG1oTm8YIZj1IuG2umya2Rg8J0sAl36nFu6YpnI+yQl/TdOJK9+3R\r
+       Xz9uDmaB/QbggCiLiN+n2kd6OSvJ9c0mfoJL7YCHipW9T4mKuAWmcx3zg5MiB2fhaZfr\r
+       mbGMms5oVoY9+xDh8oWEdMD8VKApBBD7RmmghG3F4bVmLNnTVR/6uTfQXbdbYE4fjTRA\r
+       +E3BVE16aFo+AAfVlUBRg5rl9t6HQHb4xmCy9paxhMdHLwBuOtlNJVcD1M9jPREB4HPY\r
+       5H84BlTUinzBGrETQ56Q8I10R+vmypPJYUvvYF9Qktrv8TjE18XE6aSpPSNAlssnnVze\r
+       6Gcw==\r
+MIME-Version: 1.0\r
+Received: by 10.52.17.207 with SMTP id q15mr3929265vdd.49.1341158528204; Sun,\r
+       01 Jul 2012 09:02:08 -0700 (PDT)\r
+Received: by 10.220.6.3 with HTTP; Sun, 1 Jul 2012 09:02:08 -0700 (PDT)\r
+In-Reply-To: <87k3yrmahu.fsf@qmul.ac.uk>\r
+References: <1340656899-5644-1-git-send-email-ethan@betacantrips.com>\r
+       <877gutnmf1.fsf@qmul.ac.uk>\r
+       <CAOJ+Ob0Kw0Kkhh9C27Xv9gvqtNowzQiNqrLAtvti7fL8NND2+w@mail.gmail.com>\r
+       <87k3yrmahu.fsf@qmul.ac.uk>\r
+Date: Sun, 1 Jul 2012 12:02:08 -0400\r
+Message-ID:\r
+ <CAOJ+Ob0MSOez2MvD2fCgF7t32kFPk4g2+xCud88QmBLt_b5pOA@mail.gmail.com>\r
+Subject: Re: [RFC PATCH 00/14] modular mail stores based on URIs\r
+From: Ethan <ethan.glasser.camp@gmail.com>\r
+To: Mark Walters <markwalters1009@gmail.com>\r
+Content-Type: multipart/alternative; boundary=bcaec5040a4ca93e8104c3c6cd04\r
+Cc: notmuch@notmuchmail.org\r
+X-BeenThere: notmuch@notmuchmail.org\r
+X-Mailman-Version: 2.1.13\r
+Precedence: list\r
+List-Id: "Use and development of the notmuch mail system."\r
+       <notmuch.notmuchmail.org>\r
+List-Unsubscribe: <http://notmuchmail.org/mailman/options/notmuch>,\r
+       <mailto:notmuch-request@notmuchmail.org?subject=unsubscribe>\r
+List-Archive: <http://notmuchmail.org/pipermail/notmuch>\r
+List-Post: <mailto:notmuch@notmuchmail.org>\r
+List-Help: <mailto:notmuch-request@notmuchmail.org?subject=help>\r
+List-Subscribe: <http://notmuchmail.org/mailman/listinfo/notmuch>,\r
+       <mailto:notmuch-request@notmuchmail.org?subject=subscribe>\r
+X-List-Received-Date: Sun, 01 Jul 2012 16:02:12 -0000\r
+\r
+--bcaec5040a4ca93e8104c3c6cd04\r
+Content-Type: text/plain; charset=ISO-8859-1\r
+\r
+Thanks for going through it, I know there's a lot to go through..\r
+\r
+On Thu, Jun 28, 2012 at 4:45 PM, Mark Walters <markwalters1009@gmail.com>wrote:\r
+\r
+> I was thinking of just having one mail root and inside that there could\r
+> be maildirs and mboxes. Everything would still be relative to the root.\r
+>\r
+\r
+I'm hesitant to have directories that contain maildirs and mboxes. It\r
+should be possible to unambiguously distinguish between a maildir file and\r
+an mbox file (mboxes always start with "From ", no colon) but it sounds\r
+kind of fragile.\r
+\r
+>  1. Are URIs the way to specify individual messages, despite bremner's\r
+> >  concerns about too much of the API being strings? Is adding another\r
+> library\r
+> >  is the easiest way to parse URIs?\r
+>\r
+> In my opinion  the nice thing about using strings is that it does not\r
+> require\r
+> any changes to the Xapian database to store them. I think using URIs may\r
+> not be best though as they seem to be annoying to parse (as filenames\r
+> can contain the same characters) and you seem to need to work around the\r
+> parser in some cases.\r
+>\r
+\r
+I think that's more the fault of the parser than of the URIs. If glib came\r
+with a parser, that would be great. There aren't a lot of options for\r
+pure-C URI parsing. Besides uriparser, there's also some code in the W3C\r
+sample code library, but it looked like integrating it would be a pain so I\r
+let it go.\r
+\r
+I wonder if the following would be practical: use // as the field\r
+> separator:\r
+>\r
+> e.g. mbox://filename//start_of_message+length\r
+>\r
+> I think 2 consecutive slashes // is about the only thing we can assume\r
+> is not in the path or filename. Since it is not in the filename I think\r
+> parsing should be trivial (thus avoiding the extra library).\r
+>\r
+\r
+Can you explain what you mean when you say that two consecutive slashes\r
+can't appear in a URL? Ordinary filesystem paths can contain them, and so\r
+can file: URLs. (I just looked up file:///home/ethan///////tmp and Firefox\r
+handled that OK.) I've sometimes seen machine-generated filenames with\r
+double slashes because that way you don't have to make sure the incoming\r
+filename was correctly terminated before adding another level.\r
+\r
+\r
+> Secondly, I would prefer to keep maildirs as just the bare file name: so\r
+> the existence of // can be the signal that there is some other\r
+> scheme. This is asymmetric, but is rather more backwardly compatible.\r
+>\r
+\r
+Based on your and Jani's reasoning, I did this. Revised patch series\r
+follows.\r
+\r
+Ethan\r
+\r
+--bcaec5040a4ca93e8104c3c6cd04\r
+Content-Type: text/html; charset=ISO-8859-1\r
+Content-Transfer-Encoding: quoted-printable\r
+\r
+Thanks for going through it, I know there&#39;s a lot to go through..<br><b=\r
+r><div class=3D"gmail_quote">On Thu, Jun 28, 2012 at 4:45 PM, Mark Walters =\r
+<span dir=3D"ltr">&lt;<a href=3D"mailto:markwalters1009@gmail.com" target=\r
+=3D"_blank">markwalters1009@gmail.com</a>&gt;</span> wrote:<br>\r
+\r
+<blockquote class=3D"gmail_quote" style=3D"margin:0 0 0 .8ex;border-left:1p=\r
+x #ccc solid;padding-left:1ex">I was thinking of just having one mail root =\r
+and inside that there could<br>\r
+be maildirs and mboxes. Everything would still be relative to the root.<br>=\r
+</blockquote><div><br>I&#39;m hesitant to have directories that contain mai=\r
+ldirs and mboxes. It should be possible to unambiguously distinguish betwee=\r
+n a maildir file and an mbox file (mboxes always start with &quot;From &quo=\r
+t;, no colon) but it sounds kind of fragile.<br>\r
+\r
+\r
+<br></div><blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8=\r
+ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"><div>\r
+&gt; =A01. Are URIs the way to specify individual messages, despite bremner=\r
+&#39;s<br>\r
+&gt; =A0concerns about too much of the API being strings? Is adding another=\r
+ library<br>\r
+&gt; =A0is the easiest way to parse URIs?<br>\r
+\r
+</div><br>In my opinion =A0the nice thing about using strings is that it do=\r
+es not require<br>\r
+any changes to the Xapian database to store them. I think using URIs may<br=\r
+>\r
+not be best though as they seem to be annoying to parse (as filenames<br>\r
+can contain the same characters) and you seem to need to work around the<br=\r
+>\r
+parser in some cases.<br></blockquote><div><br>I think that&#39;s more the =\r
+fault of the parser than of the URIs. If glib came with a parser, that woul=\r
+d be great. There aren&#39;t a lot of options for pure-C URI parsing. Besid=\r
+es uriparser, there&#39;s also some code in the W3C sample code library, bu=\r
+t it looked like integrating it would be a pain so I let it go.<br>\r
+\r
+<br></div><blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8=\r
+ex;border-left:1px solid rgb(204,204,204);padding-left:1ex">\r
+\r
+I wonder if the following would be practical: use // as the field<br>\r
+separator:<br>\r
+<br>\r
+e.g. mbox://filename//start_of_message+length<br>\r
+<br>\r
+I think 2 consecutive slashes // is about the only thing we can assume<br>\r
+is not in the path or filename. Since it is not in the filename I think<br>\r
+parsing should be trivial (thus avoiding the extra library).<br></blockquot=\r
+e><div><br>Can you explain what you mean when you say that two consecutive =\r
+slashes can&#39;t appear in a URL? Ordinary filesystem paths can contain th=\r
+em, and so can file: URLs. (I just looked up file:///home/ethan///////tmp a=\r
+nd Firefox handled that OK.) I&#39;ve sometimes seen  machine-generated fil=\r
+enames with double slashes because that way you don&#39;t have to make sure=\r
+ the incoming filename was correctly terminated before adding another level=\r
+.<br>\r
+\r
+=A0</div><blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8e=\r
+x;border-left:1px solid rgb(204,204,204);padding-left:1ex">\r
+\r
+Secondly, I would prefer to keep maildirs as just the bare file name: so<br=\r
+>\r
+the existence of // can be the signal that there is some other<br>\r
+scheme. This is asymmetric, but is rather more backwardly compatible.<br></=\r
+blockquote><div><br>Based on your and Jani&#39;s reasoning, I did this. Rev=\r
+ised patch series follows.<br>\r
+<br></div><div>Ethan<br><br></div></div>\r
+\r
+--bcaec5040a4ca93e8104c3c6cd04--\r