--- /dev/null
+Return-Path: <mpn@google.com>\r
+X-Original-To: notmuch@notmuchmail.org\r
+Delivered-To: notmuch@notmuchmail.org\r
+Received: from localhost (localhost [127.0.0.1])\r
+ by olra.theworths.org (Postfix) with ESMTP id 68D93431FB6\r
+ for <notmuch@notmuchmail.org>; Tue, 4 Sep 2012 12:44:04 -0700 (PDT)\r
+X-Virus-Scanned: Debian amavisd-new at olra.theworths.org\r
+X-Spam-Flag: NO\r
+X-Spam-Score: -0.7\r
+X-Spam-Level: \r
+X-Spam-Status: No, score=-0.7 tagged_above=-999 required=5\r
+ tests=[DKIM_SIGNED=0.1, DKIM_VALID=-0.1, RCVD_IN_DNSWL_LOW=-0.7]\r
+ autolearn=disabled\r
+Received: from olra.theworths.org ([127.0.0.1])\r
+ by localhost (olra.theworths.org [127.0.0.1]) (amavisd-new, port 10024)\r
+ with ESMTP id 2ysK6ajrBbF0 for <notmuch@notmuchmail.org>;\r
+ Tue, 4 Sep 2012 12:44:03 -0700 (PDT)\r
+Received: from mail-ee0-f53.google.com (mail-ee0-f53.google.com\r
+ [74.125.83.53]) (using TLSv1 with cipher RC4-SHA (128/128 bits)) (No client\r
+ certificate requested) by olra.theworths.org (Postfix) with ESMTPS id\r
+ CD1AA431FAF for <notmuch@notmuchmail.org>; Tue, 4 Sep 2012 12:44:02 -0700\r
+ (PDT)\r
+Received: by eekb47 with SMTP id b47so2969538eek.26\r
+ for <notmuch@notmuchmail.org>; Tue, 04 Sep 2012 12:44:01 -0700 (PDT)\r
+DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com;\r
+ s=20120113; h=sender:from:to:subject:in-reply-to:organization:references\r
+ :user-agent:x-face:face:x-pgp:x-pgp-fp:date:message-id:mime-version\r
+ :content-type; bh=k3cSusr4oXo9pUf0EKe6avU/BtoWgLojYDKKQ1dCp/E=;\r
+ b=FbO3vYN4Wzgm/NDjOLlRLs3RwqNcWLVu/dHcnt43tBXCyk1oud8BrIRi3z41Yq9bYX\r
+ ee5/Ekq9tybfaLtzMOt6snW9H7qI+WEmK7PMOFuA3IQdt0REsb+bjNN5SxAxbvo46yec\r
+ M3rWauDcweoV9gB7WrvU+ElKHLpIsflfY408+LY/838DEcp2pIwAquL818pxuAs2RB7g\r
+ CE021F2BJ5AdkKKwZJICAr0ViNSl8l8N+5hTm5hT7iFsSx0Eu5qj05XOiycy3h4I75tx\r
+ adiFLOf12io527H0ZAN0nyyjXMW4OuE6fY0JSg15VGnlC5BIIQGRFHgt/EwRi0xnnlfu 93jA==\r
+X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;\r
+ d=google.com; s=20120113;\r
+ h=sender:from:to:subject:in-reply-to:organization:references\r
+ :user-agent:x-face:face:x-pgp:x-pgp-fp:date:message-id:mime-version\r
+ :content-type:x-gm-message-state;\r
+ bh=k3cSusr4oXo9pUf0EKe6avU/BtoWgLojYDKKQ1dCp/E=;\r
+ b=QR50DENAyjo5YeT44qbzJR9oQkOfOQJlLoXLL8pciUICJ6lgUEaF0kq+Gv4CmmOPC0\r
+ +cgH8N/zjS5gas29+UiqxGW5YHnY/JFiTBNw9HK+tduO/dlKZihJgiwTwTq6NH0sZ5Lg\r
+ 4PyWK0d21OTQ3IoTZp6Ckm0hYewPydhv9GSukrgy6qbD2YDcIfLIRXrXqXShME17OR+M\r
+ Izlff+IYp9BFPzu+tK1rpq+dyRkIz0dybTkRwZRxi+X0YKSw7b9wvnLnWlDSa6gUQxg4\r
+ gsDpmIebmWl7of+tYE1HlqZSdlzDeYTyUfJXXZoZ4o29R9tpElh3Te4KP+gYtZ2FvLsp\r
+ zORg==\r
+Received: by 10.14.172.193 with SMTP id t41mr27637811eel.25.1346787841727;\r
+ Tue, 04 Sep 2012 12:44:01 -0700 (PDT)\r
+Received: by 10.14.172.193 with SMTP id t41mr27637799eel.25.1346787841546;\r
+ Tue, 04 Sep 2012 12:44:01 -0700 (PDT)\r
+Received: from mpn-glaptop ([2620:0:105f:5:f2de:f1ff:fe35:1a72])\r
+ by mx.google.com with ESMTPS id v3sm47922341eep.10.2012.09.04.12.43.59\r
+ (version=TLSv1/SSLv3 cipher=OTHER);\r
+ Tue, 04 Sep 2012 12:44:00 -0700 (PDT)\r
+Sender: Michal Nazarewicz <mpn@google.com>\r
+From: Michal Nazarewicz <mina86@mina86.com>\r
+To: Dmitry Kurochkin <dmitry.kurochkin@gmail.com>, notmuch@notmuchmail.org\r
+Subject: Re: [PATCH] Add notmuch-remove-duplicates.py script to contrib.\r
+In-Reply-To: <1346784785-19746-1-git-send-email-dmitry.kurochkin@gmail.com>\r
+Organization: http://mina86.com/\r
+References: <1346784785-19746-1-git-send-email-dmitry.kurochkin@gmail.com>\r
+User-Agent: Notmuch/0.14+2~g416b120 (http://notmuchmail.org) Emacs/24.2.50.1\r
+ (x86_64-unknown-linux-gnu)\r
+X-Face: PbkBB1w#)bOqd`iCe"Ds{e+!C7`pkC9a|f)Qo^BMQvy\q5x3?vDQJeN(DS?|-^$uMti[3D*#^_Ts"pU$jBQLq~Ud6iNwAw_r_o_4]|JO?]}P_}Nc&"p#D(ZgUb4uCNPe7~a[DbPG0T~!&c.y$Ur,=N4RT>]dNpd; KFrfMCylc}gc??'U2j,!8%xdD\r
+Face: iVBORw0KGgoAAAANSUhEUgAAADAAAAAwBAMAAAClLOS0AAAAJFBMVEWbfGlUPDDHgE57V0jUupKjgIObY0PLrom9mH4dFRK4gmjPs41MxjOgAAACQElEQVQ4jW3TMWvbQBQHcBk1xE6WyALX1069oZBMlq+ouUwpEQQ6uRjttkWP4CmBgGM0BQLBdPFZYPsyFUo6uEtKDQ7oy/U96XR2Ux8ehH/89Z6enqxBcS7Lg81jmSuujrfCZcLI/TYYvbGj+jbgFpHJ/bqQAUISj8iLyu4LuFHJTosxsucO4jSDNE0Hq3hwK/ceQ5sx97b8LcUDsILfk+ovHkOIsMbBfg43VuQ5Ln9YAGCkUdKJoXR9EclFBhixy3EGVz1K6eEkhxCAkeMMnqoAhAKwhoUJkDrCqvbecaYINlFKSRS1i12VKH1XpUd4qxL876EkMcDvHj3s5RBajHHMlA5iK32e0C7VgG0RlzFPvoYHZLRmAC0BmNcBruhkE0KsMsbEc62ZwUJDxWUdMsMhVqovoT96i/DnX/ASvz/6hbCabELLk/6FF/8PNpPCGqcZTGFcBhhAaZZDbQPaAB3+KrWWy2XgbYDNIinkdWAFcCpraDE/knwe5DBqGmgzESl1p2E4MWAz0VUPgYYzmfWb9yS4vCvgsxJriNTHoIBz5YteBvg+VGISQWUqhMiByPIPpygeDBE6elD973xWwKkEiHZAHKjhuPsFnBuArrzxtakRcISv+XMIPl4aGBUJm8Emk7qBYU8IlgNEIpiJhk/No24jHwkKTFHDWfPniR4iw5vJaw2nzSjfq2zffcE/GDjRC2dn0J0XwPAbDL84TvaFCJEU4Oml9pRyEUhR3Cl2t01AoEjRbs0sYugp14/4X5n4pU4EHHnMAAAAAElFTkSuQmCC\r
+X-PGP: 50751FF4\r
+X-PGP-FP: AC1F 5F5C D418 88F8 CC84 5858 2060 4012 5075 1FF4\r
+Date: Tue, 04 Sep 2012 21:43:53 +0200\r
+Message-ID: <xa1tligpk1za.fsf@mina86.com>\r
+MIME-Version: 1.0\r
+Content-Type: multipart/mixed; boundary="=-=-="\r
+X-Gm-Message-State: ALoCoQlVnkhgHJ0LZZPUpTWbE6G/J1gRct7k8MTReEQQx8SMsxxf8DVkw91sbuO04tPVb/dEgT2RYYt7cU2/CgwxISsOmNdZstU1qExzkrAiFdK3foWxXFdhYyYc4OdJoXmZ5744jTqBhqYU1dc2cBIRxLmdcu2ahUFQuIgKopDTwfAyfGsERq0GsHCaR0q8ha1GlZR6/kzGFvNu2JnzkVtnvHHgz2oWlA==\r
+X-BeenThere: notmuch@notmuchmail.org\r
+X-Mailman-Version: 2.1.13\r
+Precedence: list\r
+List-Id: "Use and development of the notmuch mail system."\r
+ <notmuch.notmuchmail.org>\r
+List-Unsubscribe: <http://notmuchmail.org/mailman/options/notmuch>,\r
+ <mailto:notmuch-request@notmuchmail.org?subject=unsubscribe>\r
+List-Archive: <http://notmuchmail.org/pipermail/notmuch>\r
+List-Post: <mailto:notmuch@notmuchmail.org>\r
+List-Help: <mailto:notmuch-request@notmuchmail.org?subject=help>\r
+List-Subscribe: <http://notmuchmail.org/mailman/listinfo/notmuch>,\r
+ <mailto:notmuch-request@notmuchmail.org?subject=subscribe>\r
+X-List-Received-Date: Tue, 04 Sep 2012 19:44:04 -0000\r
+\r
+--=-=-=\r
+Content-Type: text/plain; charset=utf-8\r
+Content-Transfer-Encoding: quoted-printable\r
+\r
+On Tue, Sep 04 2012, Dmitry Kurochkin wrote:\r
+> The script removes duplicate message files. It takes no options.\r
+>\r
+> Files are assumed duplicates if their content is the same except for\r
+> ignored headers. Currently, the only ignored header is Received:.\r
+> ---\r
+> contrib/notmuch-remove-duplicates.py | 95 ++++++++++++++++++++++++++++=\r
+++++++\r
+> 1 file changed, 95 insertions(+)\r
+> create mode 100755 contrib/notmuch-remove-duplicates.py\r
+>\r
+> diff --git a/contrib/notmuch-remove-duplicates.py b/contrib/notmuch-remov=\r
+e-duplicates.py\r
+> new file mode 100755\r
+> index 0000000..dbe2e25\r
+> --- /dev/null\r
+> +++ b/contrib/notmuch-remove-duplicates.py\r
+> @@ -0,0 +1,95 @@\r
+> +#!/usr/bin/env python\r
+> +\r
+> +import sys\r
+> +\r
+> +IGNORED_HEADERS =3D [ "Received:" ]\r
+> +\r
+> +if len(sys.argv) !=3D 1:\r
+> + print "Usage: %s" % sys.argv[0]\r
+> + print\r
+> + print "The script removes duplicate message files. Takes no options=\r
+."\r
+> + print "Requires notmuch python module."\r
+> + print\r
+> + print "Files are assumed duplicates if their content is the same"\r
+> + print "except for the following headers: %s." % ", ".join(IGNORED_HE=\r
+ADERS)\r
+> + exit(1)\r
+\r
+It's much better put inside a main() function, which is than called only\r
+if the script is run directly.\r
+\r
+> +\r
+> +import notmuch\r
+> +import os\r
+> +import time\r
+> +\r
+> +class MailComparator:\r
+> + """Checks if mail files are duplicates."""\r
+> + def __init__(self, filename):\r
+> + self.filename =3D filename\r
+> + self.mail =3D self.readFile(self.filename)\r
+> +\r
+> + def isDuplicate(self, filename):\r
+> + return self.mail =3D=3D self.readFile(filename)\r
+> +\r
+> + @staticmethod\r
+> + def readFile(filename):\r
+> + with open(filename) as f:\r
+> + data =3D ""\r
+> + while True:\r
+> + line =3D f.readline()\r
+> + for header in IGNORED_HEADERS:\r
+> + if line.startswith(header):\r
+\r
+Case of headers should be ignored, but this does not ignore it.\r
+\r
+> + # skip header continuation lines\r
+> + while True:\r
+> + line =3D f.readline()\r
+> + if len(line) =3D=3D 0 or line[0] not in [" "=\r
+, "\t"]:\r
+> + break\r
+> + break\r
+\r
+This will ignore line just after the ignored header.\r
+\r
+> + else:\r
+> + data +=3D line\r
+> + if line =3D=3D "\n":\r
+> + break\r
+> + data +=3D f.read()\r
+> + return data\r
+> +\r
+> +db =3D notmuch.Database()\r
+> +query =3D db.create_query('*')\r
+> +print "Number of messages: %s" % query.count_messages()\r
+> +\r
+> +files_count =3D 0\r
+> +for root, dirs, files in os.walk(db.get_path()):\r
+> + if not root.startswith(os.path.join(db.get_path(), ".notmuch/")):\r
+> + files_count +=3D len(files)\r
+> +print "Number of files: %s" % files_count\r
+> +print "Estimated number of duplicates: %s" % (files_count - query.count_=\r
+messages())\r
+> +\r
+> +msgs =3D query.search_messages()\r
+> +msg_count =3D 0\r
+> +suspected_duplicates_count =3D 0\r
+> +duplicates_count =3D 0\r
+> +timestamp =3D time.time()\r
+> +for msg in msgs:\r
+> + msg_count +=3D 1\r
+> + if len(msg.get_filenames()) > 1:\r
+> + filenames =3D msg.get_filenames()\r
+> + comparator =3D MailComparator(filenames.next())\r
+> + for filename in filenames:\r
+\r
+Strictly speaking, you need to compare each file to each file, and not\r
+just every file to the first file.\r
+\r
+> + if os.path.realpath(comparator.filename) =3D=3D os.path.real=\r
+path(filename):\r
+> + print "Message '%s' has filenames pointing to the\r
+> same file: '%s' '%s'" % (msg.get_message_id(), comparator.filename,\r
+> filename)\r
+\r
+So why aren't those removed?\r
+\r
+> + elif comparator.isDuplicate(filename):\r
+> + os.remove(filename)\r
+> + duplicates_count +=3D 1\r
+> + else:\r
+> + #print "Potential duplicates: %s" % msg.get_message_id()\r
+> + suspected_duplicates_count +=3D 1\r
+> +\r
+> + new_timestamp =3D time.time()\r
+> + if new_timestamp - timestamp > 1:\r
+> + timestamp =3D new_timestamp\r
+> + sys.stdout.write("\rProcessed %s messages, removed %s duplicates=\r
+..." % (msg_count, duplicates_count))\r
+> + sys.stdout.flush()\r
+> +\r
+> +print "\rFinished. Processed %s messages, removed %s duplicates." % (msg=\r
+_count, duplicates_count)\r
+> +if duplicates_count > 0:\r
+> + print "You might want to run 'notmuch new' now."\r
+> +\r
+> +if suspected_duplicates_count > 0:\r
+> + print\r
+> + print "Found %s messages with duplicate IDs but different content." =\r
+% suspected_duplicates_count\r
+> + print "Perhaps we should ignore more headers."\r
+\r
+Please consider the following instead (not tested):\r
+\r
+\r
+#!/usr/bin/env python\r
+\r
+import collections\r
+import notmuch\r
+import os\r
+import re\r
+import sys\r
+import time\r
+\r
+\r
+IGNORED_HEADERS =3D [ 'Received' ]\r
+\r
+\r
+isIgnoredHeadersLine =3D re.compile(\r
+ r'^(?:%s)\s*:' % '|'.join(IGNORED_HEADERS),\r
+ re.IGNORECASE).search\r
+\r
+doesStartWithWS =3D re.compile(r'^\s').search\r
+\r
+\r
+def usage(argv0):\r
+ print """Usage: %s [<query-string>]\r
+\r
+The script removes duplicate message files. Takes no options."\r
+Requires notmuch python module."\r
+\r
+Files are assumed duplicates if their content is the same"\r
+except for the following headers: %s.""" % (argv0, ', '.join(IGNORED_HEADER=\r
+S))\r
+\r
+\r
+def readMailFile(filename):\r
+ with open(filename) as fd:\r
+ data =3D []\r
+ skip_header =3D False\r
+ for line in fd:\r
+ if doesStartWithWS(line):\r
+ if not skip_header:\r
+ data.append(line)\r
+ elif isIgnoredHeadersLine(line):\r
+ skip_header =3D True\r
+ else:\r
+ data.append(line)\r
+ if line =3D=3D '\n':\r
+ break\r
+ data.append(fd.read())\r
+ return ''.join(data)\r
+\r
+\r
+def dedupMessage(msg):\r
+ filenames =3D msg.get_filenames()\r
+ if len(filenames) <=3D 1:\r
+ return (0, 0)\r
+\r
+ realpaths =3D collections.defaultdict(list)\r
+ contents =3D collections.defaultdict(list)\r
+ for filename in filenames:\r
+ real =3D os.path.realpath(filename)\r
+ lst =3D realpaths[real]\r
+ lst.append(filename)\r
+ if len(lst) =3D=3D 1:\r
+ contents[readMailFile(real)].append(real)\r
+\r
+ duplicates =3D 0\r
+\r
+ for filenames in contents.itervalues():\r
+ if len(filenames) > 1:\r
+ print 'Files with the same content:'\r
+ print ' ', filenames.pop()\r
+ duplicates +=3D len(filenames)\r
+ for filename in filenames:\r
+ del realpaths[filename]\r
+ # os.remane(filename)\r
+\r
+ for real, filenames in realpaths.iteritems():\r
+ if len(filenames) > 1:\r
+ print 'Files pointing to the same message:'\r
+ print ' ', filenames.pop()\r
+ duplicates +=3D len(filenames)\r
+ # for filename in filenames:\r
+ # os.remane(filename)\r
+\r
+ return (duplicates, len(realpaths) - 1)\r
+\r
+\r
+def dedupQuery(query):\r
+ print 'Number of messages: %s' % query.count_messages()\r
+ msg_count =3D 0\r
+ suspected_count =3D 0\r
+ duplicates_count =3D 0\r
+ timestamp =3D time.time()\r
+ msgs =3D query.search_messages()\r
+ for msg in msgs:\r
+ msg_count +=3D 1\r
+ d, s =3D dedupMessage(msg)\r
+ duplicates_count +=3D d\r
+ suspected_count +=3D d\r
+\r
+ new_timestamp =3D time.time()\r
+ if new_timestamp - timestamp > 1:\r
+ timestamp =3D new_timestamp\r
+ sys.stdout.write('\rProcessed %s messages, removed %s duplicate=\r
+s...'\r
+ % (msg_count, duplicates_count))\r
+ sys.stdout.flush()\r
+\r
+ print '\rFinished. Processed %s messages, removed %s duplicates.' % (\r
+ msg_count, duplicates_count)\r
+ if duplicates_count > 0:\r
+ print 'You might want to run "notmuch new" now.'\r
+\r
+ if suspected_duplicates_count > 0:\r
+ print """\r
+Found %d messages with duplicate IDs but different content.\r
+Perhaps we should ignore more headers.""" % suspected_count\r
+\r
+\r
+def main(argv):\r
+ if len(argv) =3D=3D 1:\r
+ query =3D '*'\r
+ elif len(argv) =3D=3D 2:\r
+ query =3D argv[1]\r
+ else:\r
+ usage(argv[0])\r
+ return 1\r
+\r
+ db =3D notmuch.Database()\r
+ query =3D db.create_query(query)\r
+ dedupQuery(db, query)\r
+ return 0\r
+\r
+\r
+if __name__ =3D=3D '__main__':\r
+ sys.exit(main(sys.argv))\r
+\r
+\r
+\r
+--=20\r
+Best regards, _ _\r
+.o. | Liege of Serenely Enlightened Majesty of o' \,=3D./ `o\r
+..o | Computer Science, Micha=C5=82 =E2=80=9Cmina86=E2=80=9D Nazarewicz =\r
+ (o o)\r
+ooo +----<email/xmpp: mpn@google.com>--------------ooO--(_)--Ooo--\r
+--=-=-=\r
+Content-Type: multipart/signed; boundary="==-=-=";\r
+ micalg=pgp-sha1; protocol="application/pgp-signature"\r
+\r
+--==-=-=\r
+Content-Type: text/plain\r
+\r
+\r
+--==-=-=\r
+Content-Type: application/pgp-signature\r
+\r
+-----BEGIN PGP SIGNATURE-----\r
+Version: GnuPG v1.4.10 (GNU/Linux)\r
+\r
+iQIcBAEBAgAGBQJQRln5AAoJECBgQBJQdR/0EQQP/17AJmk0zYfSsuIz4I7H3Ykm\r
+9YMt/yK1hn/u6yDtylSbnjTgWT0t6OWbOplIzmW9q6iqwhMcdqP20HXMEkqSbNLa\r
+7WJ+4VXKRLb6PC3bQmM8eqYNtflgEAEWyAuZSQv9f893e6vH/e+7yoFtxaUypcUW\r
+Wf8qm3T1ljle+2S1xGbteoVDUSHY5epesXlWR6hVA9Qclc/4xpVLNapx3EKRkxBh\r
+vOpe+u5ATa04DYvIOoGVl723PBIHpm25cGen5lc8vOjXKwqhG0G7di5E29BhAyVT\r
+yZorKrfsBRTTIlYEErakrzGhiMP3zRnCQmFWvIj/ASbiOUnX8ktFMjfqe+DNW3zq\r
+T/2jpdzhBdVyioLhBIsMGLdsW6yIk3LURcw4uTijEG2ITj9kdQspydGFahTJk1Ly\r
+cIls19AMCK7xfGBt8o3xYMX6v/bOxpz/Hot0e+SdHQtiByIUKfJMF7gMo6YyxRfh\r
+cq1mgoLm+L4/zdrf4IMZDUpoMM8q4yr3eJibINlLxAmRnD3CpnVE2wf6mExLnYxy\r
+PTIQ9p3pRHsxbRuHvYylJfNNlGpjsRFSgKeRF50iFY+TnzUh+40Tp18BTbL9Dd7R\r
+UGuIMScxZ6qKt5MQhfBw1F+JpaaIsLTMSh1MjdzCvNVVRvkx7MQuUxiEcOj7wyEf\r
+0zip9GVi2hAsumsgL0Wz\r
+=mXjE\r
+-----END PGP SIGNATURE-----\r
+--==-=-=--\r
+\r
+--=-=-=--\r