Help: cannot OCR pdf file in Arabic
Hi to everybody
I am not new to Zotervo but very new to coding.
Yesterday I followed step by step the procedure to install the OCR plugin in Zotero. I have also followed the instruction of adding the various languages of my texts, one of which is Arabic. I re-run the script today (sudo port install tesseract-ara), to be sure to have done it correctly.
Even though I can run the script apparently correctly, the OCR in arabic does not work (thoug it works in English). When I copy the text in the pdf, I see while pasting that the computer has not detected that it is arabic, and I get a combination of characters such as : ey 2l ae ;-‘u_.-i':\} ‘L;JJ}L-H
Note that the OCRisation of texts in English worked from the start quite well (even though not as powerful as preview, especially if the characters are messy).
Thans for the help.
I am not new to Zotervo but very new to coding.
Yesterday I followed step by step the procedure to install the OCR plugin in Zotero. I have also followed the instruction of adding the various languages of my texts, one of which is Arabic. I re-run the script today (sudo port install tesseract-ara), to be sure to have done it correctly.
Even though I can run the script apparently correctly, the OCR in arabic does not work (thoug it works in English). When I copy the text in the pdf, I see while pasting that the computer has not detected that it is arabic, and I get a combination of characters such as : ey 2l ae ;-‘u_.-i':\} ‘L;JJ}L-H
Note that the OCRisation of texts in English worked from the start quite well (even though not as powerful as preview, especially if the characters are messy).
Thans for the help.
Upgrade Storage
I have just found out how to do, but it does not seem to have any effect, I still see:
ey 2l ae ;-‘u_.-i':\} ‘L;JJ}L-H
Thanks for the ongoing help!
https://s3.amazonaws.com/zotero.org/images/forums/u1908110/q72gbiq6oq038b0099ja.png
Now since your settings look OK but this doesn't seem to work for you, can you share a link to the PDF, or perhaps just 2 pages? With that we can take a closer look.
Although I like Arabic very much, I am more comfortable having my library in English. Does it mean that when I want to ocr-ize an Arabic text I have to switch to Zotero in Arabic all the time?
The second issue, unrelated, is that the quality of the Ocr-ization is quite poor, even with texts in English.
Thanks again
English OCR quality: I had not read that as an actual issue in your first post. But similarly, if you share the PDF we can take a look.
As mentioned, I switched the language to Arabic in the settings, but nothing changed. After some time, actually the day after, without me doing anything, the OCR has turned to Arabic, as you can see from the example here:
دمحم دعسروتكدلاذاتسألاهبتكامنكل »موطرخلاةنيدم نع اوبتكنوريث
However, what is incredible is that the OCR has nothing to do with the text underlined!!!
https://s3.amazonaws.com/zotero.org/images/forums/u1908110/obbt4xju3gvpnevoic6s.png
Moreover now, when I go back to an English text, and by doing so, I switch back to English in the Zotero settings, it is as if the machine was unable to see that I changed.
So it gives me:
ا م ي ه ! ب ل ا ا د 8 ع 5 ل م ع ن ل ط ع د م 1
For the following text:
https://s3.amazonaws.com/zotero.org/images/forums/u1908110/jfjchgza4ac4b8og6d2u.png
And as you see, the setting are in English:
https://s3.amazonaws.com/zotero.org/images/forums/u1908110/olmdjwtzlstc49xa2b1c.png
Maybe I did something wrong with the scripts?
Many thanks!