alpha wrote,
> It does not support automatic metadata extraction from the abstract and title page, but so far I have not really found anything that will do that consistently.
Have you found anything that will extract titles and abstacts even inconsistently?
How do you manage your PDF files?
-
alpha
- Posts: 0
- Joined: Thu Jan 01, 2004 12:00 am
How do you manage your PDF files?
CB2BIB tries to extract metadata from PDFs, etc. into BibTEX format. My limited trials with with CB2BIB and PDF working papers have not been promising - it gives no results for half of the files and inconsistent results for the others.
- cpptrader
- Posts: 0
- Joined: Thu Jan 01, 2004 12:00 am
How do you manage your PDF files?
PJ:
I don't know guys, but I have lots of files in .dvi, .ps, .djvu,
even some in .txt, .doc,.chm, and .html (there are even some nice articles).
Do they all are needed to be converted into .pdf?
(That's a bugger with .djvu files)
A little late on this one, sorry for the bump. There was a file I had in .djvu and it was the only format available. I'm pretty sure that CutePDF switched it to pdf pretty straightforward. It was a couple mos ago so forgive me if I'm mistaken. Still a handy way to switch anything from websites to .doc into pdf format.
I don't know guys, but I have lots of files in .dvi, .ps, .djvu,
even some in .txt, .doc,.chm, and .html (there are even some nice articles).
Do they all are needed to be converted into .pdf?
(That's a bugger with .djvu files)
A little late on this one, sorry for the bump. There was a file I had in .djvu and it was the only format available. I'm pretty sure that CutePDF switched it to pdf pretty straightforward. It was a couple mos ago so forgive me if I'm mistaken. Still a handy way to switch anything from websites to .doc into pdf format.
If I owe a million dollars I am lost. But if I owe $50 billion the bankers are lost. -Celso Ming
-
alpha
- Posts: 0
- Joined: Thu Jan 01, 2004 12:00 am
How do you manage your PDF files?
Zotero released an improved version this summer that adds full-text PDF indexing, so that you are no longer limited to the abstract and other structured text fields when searching.
My PDF workflow now looks like this (I do not have existing bibligraphic information in BibTex):
My PDF workflow now looks like this (I do not have existing bibligraphic information in BibTex):
- Find a PDF paper to import, and open it in a Firefox tab
- In another Firefox tab, search for the name of the PDF file on Google Scholar or another site from which Zotero can scrape the bibliographic information. A small document icon appears in Firefox (to the right of the URL) indicating that the page can be scraped. Google Scholar itself only captures paper title and author names, but if you navigate to the page of the journal publisher (ACM, IEEE, etc.) or to JSTOR, you will often be able to capture abstract, journal name, etc.
- Click the small document icon to import the bibliographic information, and on a few sites the PDF document as well
- If previous step does not grab the PDF document, switch Firefox focus to the PDF file, switch Zotero focus to the newly imported document, and then click Attachments->Add->Snapshot of Current Page. This will download the PDF to your hard drive and link it to the bibliographic information.
-
bradpreston
- Posts: 0
- Joined: Thu Jan 01, 2004 12:00 am
How do you manage your PDF files?
Not sure if its been posted here before but I have just come across http://www.mendeley.com/ and it looks really usefull. The desktop app allows you to import papers and sort into categories and it has automatic metadata extration and file renaming.
- Chuck
- Posts: 0
- Joined: Thu Jan 01, 2004 12:00 am
- AVt
- Posts: 0
- Joined: Thu Jan 01, 2004 12:00 am
How do you manage your PDF files?
Mendeley ... does it allow (indexed) text search within documents?
Something like Google Desktop (which I do not want to have on my (private) PC ...)?
Something like Google Desktop (which I do not want to have on my (private) PC ...)?
- AVt
- Posts: 0
- Joined: Thu Jan 01, 2004 12:00 am
How do you manage your PDF files?
I play with Copernic, found through Wikipedia http://en.wikipedia.org/wiki/List_of_search_engines and the benchmark study at the bottom of that link.
First at work (in case of troubles I could try for support ...) and now at home: needs a bit to configure (since I do not want media files or indexing of mail or browsing. Or too much integration into my browser).
But then it seems not bad, even previews into word or Excel files are given.
Somewhat nasty: advertising for the free version (that would drive crazy at work, but it is not free for that). But at home I tell the firewall to block web access - which seems to do what it should
Any experiences with it?
My main motivation was that I do not want Google to sniff on my PC and possibly store on the web (we had such shit in the company with internal documents).
First at work (in case of troubles I could try for support ...) and now at home: needs a bit to configure (since I do not want media files or indexing of mail or browsing. Or too much integration into my browser).
But then it seems not bad, even previews into word or Excel files are given.
Somewhat nasty: advertising for the free version (that would drive crazy at work, but it is not free for that). But at home I tell the firewall to block web access - which seems to do what it should
Any experiences with it?
My main motivation was that I do not want Google to sniff on my PC and possibly store on the web (we had such shit in the company with internal documents).
- Dimatrix
- Posts: 0
- Joined: Thu Jan 01, 2004 12:00 am
How do you manage your PDF files?
Interesting, I'll try. After the whole discussion, I still haven't found anything perfect. Uninstalled Google Desktop since it was to greedy. I wasn't happy with Mendeley. Don't remember anymore why. I deinstalled it. I'm still using JabRef for my docs, due to the bibtex tools. I'll try Copernic. All I want is something that is standard on a Mac, but still not standard on Windows. Is the searching system on Vista better?
Ctrl - L.
- AVt
- Posts: 0
- Joined: Thu Jan 01, 2004 12:00 am
How do you manage your PDF files?
Hm ... I do not find a way to extract and index from postscript or djvu files (or others), they are simply considered as text files. And the preview for those just shows them as text files (while double click starts the applications). A bit disappointing ... may be, I do not use it properly. And when I have a large list of results (say 3000 files) and click clear for the search, then it hangs up