How do you manage your PDF files?

User avatar
montpelier
Posts: 0
Joined: Thu Jan 01, 2004 12:00 am

How do you manage your PDF files?

Post by montpelier »

alpha wrote,

> It does not support automatic metadata extraction from the abstract and title page, but so far I have not really found anything that will do that consistently.



Have you found anything that will extract titles and abstacts even inconsistently?
alpha
Posts: 0
Joined: Thu Jan 01, 2004 12:00 am

How do you manage your PDF files?

Post by alpha »

CB2BIB tries to extract metadata from PDFs, etc. into BibTEX format. My limited trials with with CB2BIB and PDF working papers have not been promising - it gives no results for half of the files and inconsistent results for the others.
User avatar
cpptrader
Posts: 0
Joined: Thu Jan 01, 2004 12:00 am

How do you manage your PDF files?

Post by cpptrader »

PJ:
I don't know guys, but I have lots of files in .dvi, .ps, .djvu,

even some in .txt, .doc,.chm, and .html (there are even some nice articles).

Do they all are needed to be converted into .pdf?

(That's a bugger with .djvu files)




A little late on this one, sorry for the bump.  There was a file I had in .djvu and it was the only format available.  I'm pretty sure that CutePDF switched it to pdf pretty straightforward.  It was a couple mos ago so forgive me if I'm mistaken.  Still a handy way to switch anything from websites to .doc into pdf format.
If I owe a million dollars I am lost. But if I owe $50 billion the bankers are lost. -Celso Ming
alpha
Posts: 0
Joined: Thu Jan 01, 2004 12:00 am

How do you manage your PDF files?

Post by alpha »

Zotero released an improved version this summer that adds full-text PDF indexing, so that you are no longer limited to the abstract and other structured text fields when searching.

My PDF workflow now looks like this (I do not have existing bibligraphic information in BibTex):

  1. Find a PDF paper to import, and open it in a Firefox tab
  2. In another Firefox tab, search for the name of the PDF file on Google Scholar or another site from which Zotero can scrape the bibliographic information. A small document icon appears in Firefox (to the right of the URL) indicating that the page can be scraped. Google Scholar itself only captures paper title and author names, but if you navigate to the page of the journal publisher (ACM, IEEE, etc.) or to JSTOR, you will often be able to capture abstract, journal name, etc.
  3. Click the small document icon to import the bibliographic information, and on a few sites the PDF document as well
  4. If previous step does not grab the PDF document, switch Firefox focus to the PDF file, switch Zotero focus to the newly imported document, and then click Attachments->Add->Snapshot of Current Page. This will download the PDF to your hard drive and link it to the bibliographic information.
The entire process takes about 25 seconds, and gives you the ability to search both structured bibliographic data and the full-text PDF index.
bradpreston
Posts: 0
Joined: Thu Jan 01, 2004 12:00 am

How do you manage your PDF files?

Post by bradpreston »

Not sure if its been posted here before but I have just come across http://www.mendeley.com/ and it looks really usefull.  The desktop app allows you to import papers and sort into categories and it has automatic metadata extration and file renaming.
User avatar
Chuck
Posts: 0
Joined: Thu Jan 01, 2004 12:00 am

How do you manage your PDF files?

Post by Chuck »

Mendeley works quite well!  Thanks for posting
Speculator
User avatar
AVt
Posts: 0
Joined: Thu Jan 01, 2004 12:00 am

How do you manage your PDF files?

Post by AVt »

Mendeley ... does it allow (indexed) text search within documents?



Something like Google Desktop (which I do not want to have on my (private) PC ...)?
User avatar
AVt
Posts: 0
Joined: Thu Jan 01, 2004 12:00 am

How do you manage your PDF files?

Post by AVt »

I play with Copernic, found through Wikipedia http://en.wikipedia.org/wiki/List_of_search_engines and the benchmark study at the bottom of that link.



First at work (in case of troubles I could try for support ...) and now at home: needs a bit to configure (since I do not want media files or indexing of mail or browsing. Or too much integration into my browser).



But then it seems not bad, even previews into word or Excel files are given.



Somewhat nasty: advertising for the free version (that would drive crazy at work, but it is not free for that). But at home I tell the firewall to block web access - which seems to do what it should :-)



Any experiences with it?



My main motivation was that I do not want Google to sniff on my PC and possibly store on the web (we had such shit in the company with internal documents).
User avatar
Dimatrix
Posts: 0
Joined: Thu Jan 01, 2004 12:00 am

How do you manage your PDF files?

Post by Dimatrix »

Interesting, I'll try. After the whole discussion, I still haven't found anything perfect. Uninstalled Google Desktop since it was to greedy. I wasn't happy with Mendeley. Don't remember anymore why. I deinstalled it. I'm still using JabRef for my docs, due to the bibtex tools. I'll try Copernic. All I want is something that is standard on a Mac, but still not standard on Windows. Is the searching system on Vista better?
Ctrl - L.
User avatar
AVt
Posts: 0
Joined: Thu Jan 01, 2004 12:00 am

How do you manage your PDF files?

Post by AVt »

Hm ... I do not find a way to extract and index from postscript or djvu files (or others), they are simply considered as text files. And the preview for those just shows them as text files (while double click starts the applications). A bit disappointing ... may be, I do not use it properly. And when I have a large list of results (say 3000 files) and click clear for the search, then it hangs up
Post Reply