Page 3 of 4

How do you manage your PDF files?

Posted: Tue May 29, 2007 3:39 pm
by montpelier
alpha wrote,

> It does not support automatic metadata extraction from the abstract and title page, but so far I have not really found anything that will do that consistently.



Have you found anything that will extract titles and abstacts even inconsistently?

How do you manage your PDF files?

Posted: Tue May 29, 2007 11:01 pm
by alpha
CB2BIB tries to extract metadata from PDFs, etc. into BibTEX format. My limited trials with with CB2BIB and PDF working papers have not been promising - it gives no results for half of the files and inconsistent results for the others.

How do you manage your PDF files?

Posted: Thu Jun 28, 2007 8:49 pm
by cpptrader
PJ:
I don't know guys, but I have lots of files in .dvi, .ps, .djvu,

even some in .txt, .doc,.chm, and .html (there are even some nice articles).

Do they all are needed to be converted into .pdf?

(That's a bugger with .djvu files)




A little late on this one, sorry for the bump.  There was a file I had in .djvu and it was the only format available.  I'm pretty sure that CutePDF switched it to pdf pretty straightforward.  It was a couple mos ago so forgive me if I'm mistaken.  Still a handy way to switch anything from websites to .doc into pdf format.

How do you manage your PDF files?

Posted: Wed Oct 10, 2007 4:24 pm
by alpha
Zotero released an improved version this summer that adds full-text PDF indexing, so that you are no longer limited to the abstract and other structured text fields when searching.

My PDF workflow now looks like this (I do not have existing bibligraphic information in BibTex):

  1. Find a PDF paper to import, and open it in a Firefox tab
  2. In another Firefox tab, search for the name of the PDF file on Google Scholar or another site from which Zotero can scrape the bibliographic information. A small document icon appears in Firefox (to the right of the URL) indicating that the page can be scraped. Google Scholar itself only captures paper title and author names, but if you navigate to the page of the journal publisher (ACM, IEEE, etc.) or to JSTOR, you will often be able to capture abstract, journal name, etc.
  3. Click the small document icon to import the bibliographic information, and on a few sites the PDF document as well
  4. If previous step does not grab the PDF document, switch Firefox focus to the PDF file, switch Zotero focus to the newly imported document, and then click Attachments->Add->Snapshot of Current Page. This will download the PDF to your hard drive and link it to the bibliographic information.
The entire process takes about 25 seconds, and gives you the ability to search both structured bibliographic data and the full-text PDF index.

How do you manage your PDF files?

Posted: Sun May 03, 2009 9:27 pm
by bradpreston
Not sure if its been posted here before but I have just come across http://www.mendeley.com/ and it looks really usefull.  The desktop app allows you to import papers and sort into categories and it has automatic metadata extration and file renaming.

How do you manage your PDF files?

Posted: Sun Jun 21, 2009 7:16 pm
by Chuck
Mendeley works quite well!  Thanks for posting

How do you manage your PDF files?

Posted: Thu Jun 25, 2009 9:27 pm
by AVt
Mendeley ... does it allow (indexed) text search within documents?



Something like Google Desktop (which I do not want to have on my (private) PC ...)?

How do you manage your PDF files?

Posted: Thu Jul 02, 2009 9:24 pm
by AVt
I play with Copernic, found through Wikipedia http://en.wikipedia.org/wiki/List_of_search_engines and the benchmark study at the bottom of that link.



First at work (in case of troubles I could try for support ...) and now at home: needs a bit to configure (since I do not want media files or indexing of mail or browsing. Or too much integration into my browser).



But then it seems not bad, even previews into word or Excel files are given.



Somewhat nasty: advertising for the free version (that would drive crazy at work, but it is not free for that). But at home I tell the firewall to block web access - which seems to do what it should :-)



Any experiences with it?



My main motivation was that I do not want Google to sniff on my PC and possibly store on the web (we had such shit in the company with internal documents).

How do you manage your PDF files?

Posted: Fri Jul 03, 2009 8:27 am
by Dimatrix
Interesting, I'll try. After the whole discussion, I still haven't found anything perfect. Uninstalled Google Desktop since it was to greedy. I wasn't happy with Mendeley. Don't remember anymore why. I deinstalled it. I'm still using JabRef for my docs, due to the bibtex tools. I'll try Copernic. All I want is something that is standard on a Mac, but still not standard on Windows. Is the searching system on Vista better?

How do you manage your PDF files?

Posted: Sat Jul 04, 2009 7:38 pm
by AVt
Hm ... I do not find a way to extract and index from postscript or djvu files (or others), they are simply considered as text files. And the preview for those just shows them as text files (while double click starts the applications). A bit disappointing ... may be, I do not use it properly. And when I have a large list of results (say 3000 files) and click clear for the search, then it hangs up