Looking for some translators Latin or Greek

  • Thread starter Thread starter COPLAND_3
  • Start date Start date
Status
Not open for further replies.
C

COPLAND_3

Guest
I am thinking about starting a translating project of some of the ancient Bible commentaries of the Church Fathers, only for the purpose of translating some of the commentaries that have not yet been translated into English yet. I have all the Greek and Latin texts. I want to make as much available as possible, not for any other reason than for edification of the people of God. This is not for money or any other material gain. I will provide them on my website. If anyone is interested then please let me know. Thanks!
 
I am thinking about starting a translating project of some of the ancient Bible commentaries of the Church Fathers, only for the purpose of translating some of the commentaries that have not yet been translated into English yet. I have all the Greek and Latin texts. I want to make as much available as possible, not for any other reason than for edification of the people of God. This is not for money or any other material gain. I will provide them on my website. If anyone is interested then please let me know. Thanks!
I am interested in the Greek. What do you have and in what form are they? Are they scanned images or are they text files with Greek?
 
Dan,

litteralchristianlibrary.wetpaint.com/page/Interactive+Translating+project+of+the+Church+Fathers+Bible+Commentaries

Here are some that I have in mind so far. Right now my main resource must have a glitch, but I expect some of the others to be straightened out soon.
There might be some real undiscovered treasure if they have never been translated into English. Have they been translated into German or some other modern language? Also, Lampe would come in handy and I don’t have that lexicon in my library. Never could justify the expense.
 
Dan,

litteralchristianlibrary.wetpaint.com/page/Interactive+Translating+project+of+the+Church+Fathers+Bible+Commentaries

Here are some that I have in mind so far. Right now my main resource must have a glitch, but I expect some of the others to be straightened out soon.
Forgive me for poking around your web site 🙂 I found the list of references to the One True God very interesting.

If you look in my signature at the link for monotheism I have similar quotes from Scripture. Have you notices that in your quotes the majority of them also identify that One God of Monotheism to be the Father of Jesus?
 
Dan, to my knowledge those texts have never been translated into English, but maybe Latin. Some of them have a Greek text with a Latin translation with it, which makes it nice if you are familar with both languages. I had that luxery when I was translating Arethas’ commentary on Revelation. For some reason the link is not working right now, hopefully it will soon.

I think I may be having some others start translating soon, one is Peter Papoutsis, who is a wonderful translator, and has published some work of his own peterpapoutsis.com/
 
Hi Copland!

Several of the links do not come up, but I got a good look at Eusebius’ text on Daniel.
Within the text I am noting two separate indexing schemes: eg: the standard numbering according to the Greek alphabet: α,β,γ,δ,ε,μ,ν… etc.
And I also see the standard digitized numbers off to the left of the text for each logical sentence. What is the relative importance of each?

I have programmed my Unix system to operate on Unicode fonts both Greek, and recently Hebrew (still experimental), but it would be a bit of a pain to hand transcribe each word of the text(s) into a file to be processed the way I have gotten used to.
I can strip pdf’s of the text and make utf8 text files from them – but that will remove the accents.

Would a text file (unicode/utf-8) be acceptable as a reply format, and what is important to preserve in the texts them selves; eg: I note what appear to be computer generated analysis charts following the text itself which have no meaning to myself – but perhaps are important for others? Also, would a straight forward translation with un-accented interlinear Greek be desirable – or do you want greek and English on separate pages? (Obviously, someone could edit the accents back in later if they are desired, they really don’t affect translation in my experience, though.)

Thanks,
–Andrew.
 
Hi Copland!

Several of the links do not come up, but I got a good look at Eusebius’ text on Daniel.
Within the text I am noting two separate indexing schemes: eg: the standard numbering according to the Greek alphabet: α,β,γ,δ,ε,μ,ν… etc.
And I also see the standard digitized numbers off to the left of the text for each logical sentence. What is the relative importance of each?

I have programmed my Unix system to operate on Unicode fonts both Greek, and recently Hebrew (still experimental), but it would be a bit of a pain to hand transcribe each word of the text(s) into a file to be processed the way I have gotten used to.
I can strip pdf’s of the text and make utf8 text files from them – but that will remove the accents.

Would a text file (unicode/utf-8) be acceptable as a reply format, and what is important to preserve in the texts them selves; eg: I note what appear to be computer generated analysis charts following the text itself which have no meaning to myself – but perhaps are important for others? Also, would a straight forward translation with un-accented interlinear Greek be desirable – or do you want greek and English on separate pages? (Obviously, someone could edit the accents back in later if they are desired, they really don’t affect translation in my experience, though.)

Thanks,
–Andrew.
The format is a bit of a nuisance I admit. What I have done, which is my preference, is to use an already existing English translation for the Scripture passages, and then I translate the Greek text that the Church Fathers wrote when they were commenting on it. It does not matter to me how literal or loose one makes their translation as long as its readable. If the Septuagint is the text used by the commentator, then i would use Brenton’s English translation to save time, unless of course you want to translate that too. I used the NRSV-CE for my NT portion of the commentary. It saves lots of time doing it that way.

I plan on providing more resources soon. I am hoping that the site that I am getting them from will bring back some of what they have so-called temporially broken links. I know where to get some more texts and will provide them as well very soon.

You can look at my translation of the commentary of Philemon to how I did it. but like I said, it does not matter much to me how a person wants to lay it out. Personally I like a commentary that gives the Scripture in bold letters and then skip a line or two and have the comments underneath it, that makes it easy to follow.

Here are some helpful resources that I use that you might want to look at

Unicode writer [users.ox.ac.uk/~tayl0010/polytonic-greek-(name removed by moderator)utter.html](http://users.ox.ac.uk/~tayl0010/polytonic-greek-(name removed by moderator)utter.html)

Greek translator mymemory.translated.net/s.php?q=%CE%BA%CE%B1%CE%BA%CE%B9%CE%B1%CE%BD&sl=el-GR&tl=en-GB&sj=all

English Septuagint ecmarsh.com/lxx/

Lexicon archimedes.fas.harvard.edu/pollux/

Greek bible tools litteralchristianlibrary.wetpaint.com/page/GREEK+BIBLE+TOOLS

The Greek translator is a little helpful, but its no miracle worker to say the least. I do have some techniques that I do to help me find the meaning to words if you would like me to share them.

If you are interested in doing this, you are more than welcome to join my site as a writer and create you own page and lay it out however you want. Its all free and easy to do. Let me know if you have any questions about how to do that.

Hope to have you on board with this!
 
Hi Copland!

Several of the links do not come up, but I got a good look at Eusebius’ text on Daniel.
Within the text I am noting two separate indexing schemes: eg: the standard numbering according to the Greek alphabet: α,β,γ,δ,ε,μ,ν… etc.
And I also see the standard digitized numbers off to the left of the text for each logical sentence. What is the relative importance of each?

I have programmed my Unix system to operate on Unicode fonts both Greek, and recently Hebrew (still experimental), but it would be a bit of a pain to hand transcribe each word of the text(s) into a file to be processed the way I have gotten used to.
I can strip pdf’s of the text and make utf8 text files from them – but that will remove the accents.

Would a text file (unicode/utf-8) be acceptable as a reply format, and what is important to preserve in the texts them selves; eg: I note what appear to be computer generated analysis charts following the text itself which have no meaning to myself – but perhaps are important for others? Also, would a straight forward translation with un-accented interlinear Greek be desirable – or do you want greek and English on separate pages? (Obviously, someone could edit the accents back in later if they are desired, they really don’t affect translation in my experience, though.)

Thanks,
–Andrew.
What are you using on Unix to convert to unicode, sed or awk? I would be interested in any scripts that might be useful for converting between formats in Unix.
 
What are you using on Unix to convert to unicode, sed or awk? I would be interested in any scripts that might be useful for converting between formats in Unix.
Hi Dan!
I am a bit of a hard core programmer. I often use sed, but haven’t bothered to learn awk. Bash, sed, perl, php and python are usually enough for me. When I am in a hurry, or need to do bulk processing I ususally program in C (and rarely C++) – or even Java since the source code files of Java can include foreign characters making them easy to debug. For example, I took a transliterated hebrew file and wrote a Java program to reverse compile it back to the unicode (utf-8) characters.

Virtually all unix systems have emacs and or vim. I like vim for its easy scriptability – and I can teach it to do things like open multiple bibles and synchronize the passage being examined, executing a lexicon lookup of closest words, etc,etc.
It also quite nicely can execute and consume output from conversion scripts written in any of the aforementioned languages.

There is an iconv library, from the Gnu project, which handles most conversions between character sets if you wish to avoid doing a whole lot of scripting on your own – but to be honest, unicode (utf-8) a multibyte representation tends to be the most compatible and space efficient of the codes I have used. Most Unix command line utilities will work with it without modification, though sed and vim and Java already support unicode in its entirety, and text editors seem to be able to take unicode characters in from the keyboard and turn them into whatever format the word processors like. I have seen some people complain that “unicode services” is supposed to be better than iconv, but I have yet to have a problem worth mentioning.

I’m going to have to try the text to pdf conversion program, or perhaps Tex if that doesn’t work, to see if I can re-generate a postscript/pdf document from unicode.
That’s one of the few things I haven’t tried yet.

If there is any particular conversion you are interested in doing – let me know, and I’ll dig around a bit and see if I have a copy of something that will work already.

–Andrew.
 
Hi Dan!
I am a bit of a hard core programmer. I often use sed, but haven’t bothered to learn awk. Bash, sed, perl, php and python are usually enough for me. When I am in a hurry, or need to do bulk processing I ususally program in C (and rarely C++) – or even Java since the source code files of Java can include foreign characters making them easy to debug. For example, I took a transliterated hebrew file and wrote a Java program to reverse compile it back to the unicode (utf-8) characters.

Virtually all unix systems have emacs and or vim. I like vim for its easy scriptability – and I can teach it to do things like open multiple bibles and synchronize the passage being examined, executing a lexicon lookup of closest words, etc,etc.
It also quite nicely can execute and consume output from conversion scripts written in any of the aforementioned languages.

There is an iconv library, from the Gnu project, which handles most conversions between character sets if you wish to avoid doing a whole lot of scripting on your own – but to be honest, unicode (utf-8) a multibyte representation tends to be the most compatible and space efficient of the codes I have used. Most Unix command line utilities will work with it without modification, though sed and vim and Java already support unicode in its entirety, and text editors seem to be able to take unicode characters in from the keyboard and turn them into whatever format the word processors like. I have seen some people complain that “unicode services” is supposed to be better than iconv, but I have yet to have a problem worth mentioning.

I’m going to have to try the text to pdf conversion program, or perhaps Tex if that doesn’t work, to see if I can re-generate a postscript/pdf document from unicode.
That’s one of the few things I haven’t tried yet.

If there is any particular conversion you are interested in doing – let me know, and I’ll dig around a bit and see if I have a copy of something that will work already.

–Andrew.
I would like to convert from the Windows bibleworks fonts to unicode including the accents. BTW vi is my all time favorite editor.
 
Hey cool,
I’m no translator but please keep the forum apprised of completed texts so we can benefit.

Peterk
 
I would like to convert from the Windows bibleworks fonts to unicode including the accents. BTW vi is my all time favorite editor.
There are two separate issues here; Fonts are renderings of the underlying code – and the underlying code itself may be of several varieties. I just looked at the bible-works website fonts: bibleworks.com/fonts.html
and these are free for use with some restrictions.

I run X11 with a true type font renderer – and took my fonts off my dead apple iBook and extracted them for use under Linux. They are clean and look great and I own the copy legally to boot! (Apple logo on my IBM cut out from the lid … a trophy!)

Since I see a macintosh font package on the bibleworks site which I can extract, if you like that font particularly – I suspect I can unpack it for use on a unix machine relatively easily. Do you need help installing a font for Unix/X11 or do you already have a satisfactory one?

Digging around a bit more, the site also indicates that it exports in unicode anyway.
Do you have a sample of the text with most of the letters in it that you want to convert to utf-8; eg: the most versatile Unix variety?
If so – I can tailor one of my scripts to convert it for you; or if you don’t want to mess with installing a script – you could send me zipped packages of files you want to convert – and I can send you a zipped package of converted files with the same names; with .utf8 attached to the end so that vim will recognize the file type.

Oh – but watch out, some of these companies such as the Logos library tend to encrypt their files, and I can’t legally break encryption… but hopefully that’s not the case. (As if 400+ year old items could still be in copyright… ah well…)

Let me know what you would like to do.

–Andrew.
 
There are two separate issues here; Fonts are renderings of the underlying code – and the underlying code itself may be of several varieties. I just looked at the bible-works website fonts: bibleworks.com/fonts.html
and these are free for use with some restrictions.

I run X11 with a true type font renderer – and took my fonts off my dead apple iBook and extracted them for use under Linux. They are clean and look great and I own the copy legally to boot! (Apple logo on my IBM cut out from the lid … a trophy!)

Since I see a macintosh font package on the bibleworks site which I can extract, if you like that font particularly – I suspect I can unpack it for use on a unix machine relatively easily. Do you need help installing a font for Unix/X11 or do you already have a satisfactory one?

Digging around a bit more, the site also indicates that it exports in unicode anyway.
Do you have a sample of the text with most of the letters in it that you want to convert to utf-8; eg: the most versatile Unix variety?
If so – I can tailor one of my scripts to convert it for you; or if you don’t want to mess with installing a script – you could send me zipped packages of files you want to convert – and I can send you a zipped package of converted files with the same names; with .utf8 attached to the end so that vim will recognize the file type.

Oh – but watch out, some of these companies such as the Logos library tend to encrypt their files, and I can’t legally break encryption… but hopefully that’s not the case. (As if 400+ year old items could still be in copyright… ah well…)

Let me know what you would like to do.

–Andrew.
Thanks for the offer! I don’t run a gui linux. I currently use linux to script myself and to convert texts to html format for my own personal use. I have programmed a lot with awk, sed, grep, etc and a little with C, C++, Java and Python. In what language are your scripts written?

John 1 probably represents the entire alphabet.

WHO John 1:1 VEn avrch/| h=n o lo,goj kai. o lo,goj h=n pro.j to.n qeo,n kai. qeo.j h=n o lo,goj 2 ou-toj h=n evn avrch/| pro.j to.n qeo,n 3 pa,nta di auvtou/ evge,neto kai. cwri.j auvtou/ evge,neto ouvde. e[n o] ge,gonen 4 evn auvtw/| zwh. h=n kai. h zwh. h=n to. fw/j tw/n avnqrw,pwn\ 5 kai. to. fw/j evn th/| skoti,a| fai,nei kai. h skoti,a auvto. ouv kate,laben 6 VEge,neto a;nqrwpoj avpestalme,noj para. qeou/ o;noma auvtw/| VIwa,nnhj\ 7 ou-toj h=lqen eivj marturi,an i[na marturh,sh| peri. tou/ fwto,j i[na pa,ntej pisteu,swsin di auvtou/ 8 ouvk h=n evkei/noj to. fw/j avll i[na marturh,sh| peri. tou/ fwto,j 9 +Hn to. fw/j to. avlhqino,n o] fwti,zei pa,nta a;nqrwpon evrco,menon eivj to.n ko,smon 10 evn tw/| ko,smw| h=n kai. o ko,smoj di auvtou/ evge,neto kai. o ko,smoj auvto.n ouvk e;gnw 11 eivj ta. i;dia h=lqen kai. oi i;dioi auvto.n ouv pare,labon 12 o[soi de. e;labon auvto,n e;dwken auvtoi/j evxousi,an te,kna qeou/ gene,sqai toi/j pisteu,ousin eivj to. o;noma auvtou/ 13 oi] ouvk evx aima,twn ouvde. evk qelh,matoj sarko.j ouvde. evk qelh,matoj avndro.j avll evk qeou/ evgennh,qhsan 14 Kai. o lo,goj sa.rx evge,neto kai. evskh,nwsen evn hmi/n kai. evqeasa,meqa th.n do,xan auvtou/ do,xan wj monogenou/j para. patro,j plh,rhj ca,ritoj kai. avlhqei,aj 15 VIwa,nnhj marturei/ peri. auvtou/ kai. ke,kragen le,gwn Ou-toj h=n ~O eivpw,n ~O ovpi,sw mou evrco,menoj e;mprosqe,n mou ge,gonen o[ti prw/to,j mou h=n 16 o[ti evk tou/ plhrw,matoj auvtou/ hmei/j pa,ntej evla,bomen kai. ca,rin avnti. ca,ritoj\ 17 o[ti o no,moj dia. Mwu?se,wj evdo,qh h ca,rij kai. h avlh,qeia dia. VIhsou/ Cristou/ evge,neto 18 qeo.n ouvdei.j ew,raken pw,pote\ monogenh.j qeo.j o w’n eivj to.n ko,lpon tou/ patro.j evkei/noj evxhgh,sato 19 Kai. au[th evsti.n h marturi,a tou/ VIwa,nnou o[te avpe,steilan pro.j auvto.n oi VIoudai/oi evx ~Ierosolu,mwn ierei/j kai. Leui,taj i[na evrwth,swsin auvto,n Su. ti,j ei= 20 kai. wmolo,ghsen kai. ouvk hvrnh,sato kai. wmolo,ghsen o[ti VEgw. ouvk eivmi. o Cristo,j 21 kai. hvrw,thsan auvto,n Ti, ou=n ÎSu,Ð VHli,aj ei= kai. le,gei Ouvk eivmi, ~O profh,thj ei= su, kai. avpekri,qh Ou; 22 ei=pan ou=n auvtw/| Ti,j ei= i[na avpo,krisin dw/men toi/j pe,myasin hma/j\ ti, le,geij peri. seautou/ 23 e;fh VEgw. fwnh. bow/ntoj evn th/| evrh,mw| Euvqu,nate th.n odo.n kuri,ou kaqw.j ei=pen VHsai<aj o profh,thj 24 Kai. avpestalme,noi h=san evk tw/n Farisai,wn 25 kai. hvrw,thsan auvto.n kai. ei=pan auvtw/| Ti, ou=n bapti,zeij eiv su. ouvk ei= o Cristo.j ouvde. VHli,aj ouvde. o profh,thj 26 avpekri,qh auvtoi/j o VIwa,nnhj le,gwn VEgw. bapti,zw evn u[dati\ me,soj umw/n sth,kei o]n umei/j ouvk oi;date 27 ovpi,sw mou evrco,menoj ou- ouvk eivmi. Îevgw.Ð a;xioj i[na lu,sw auvtou/ to.n ima,nta tou/ upodh,matoj 28 Tau/ta evn Bhqani,a| evge,neto pe,ran tou/ VIorda,nou o[pou h=n o VIwa,nnhj bapti,zwn 29 Th/| evpau,rion ble,pei to.n VIhsou/n evrco,menon pro.j auvto,n kai. le,gei :Ide o avmno.j tou/ qeou/ o ai;rwn th.n amarti,an tou/ ko,smou 30 ou-to,j evstin upe.r ou- evgw. ei=pon VOpi,sw mou e;rcetai avnh.r o]j e;mprosqe,n mou ge,gonen o[ti prw/to,j mou h=n 31 kavgw. ouvk h;|dein auvto,n avll i[na fanerwqh/| tw/| VIsrah.l dia. tou/to h=lqon evgw. evn u[dati bapti,zwn 32 Kai. evmartu,rhsen VIwa,nnhj le,gwn o[ti Teqe,amai to. pneu/ma katabai/non wj peristera.n evx ouvranou/ kai. e;meinen evp auvto,n 33 kavgw. ouvk h;|dein auvto,n avll o pe,myaj me bapti,zein evn u[dati evkei/no,j moi ei=pen VEf o]n a'n i;dh|j to. pneu/ma katabai/non kai. me,non evp auvto,n ou-to,j evstin o bapti,zwn evn pneu,mati agi,w| 34 kavgw. ew,raka kai. memartu,rhka o[ti ou-to,j evstin o uio.j tou/ qeou/ 35 Th/| evpau,rion pa,lin eisth,kei VIwa,nnhj kai. evk tw/n maqhtw/n auvtou/ du,o 36 kai. evmble,yaj tw/| VIhsou/ peripatou/nti le,gei :Ide o avmno.j tou/ qeou/ 37 kai. h;kousan oi du,o maqhtai. auvtou/ lalou/ntoj kai. hvkolou,qhsan tw/| VIhsou/ 38 strafei.j de. o VIhsou/j kai. qeasa,menoj auvtou.j avkolouqou/ntaj le,gei auvtoi/j Ti, zhtei/te oi de. ei=pan auvtw/| ~Rabbi, o] le,getai meqermhneuo,menon Dida,skale pou/ me,neij 39 le,gei auvtoi/j :Ercesqe kai. o;yesqe h=lqan ou=n kai. ei=dan pou/ me,nei kai. par auvtw/| e;meinan th.n hme,ran evkei,nhn\ w[ra h=n wj deka,th 40 +Hn VAndre,aj o avdelfo.j Si,mwnoj Pe,trou ei-j evk tw/n du,o tw/n avkousa,ntwn para. VIwa,nnou kai. avkolouqhsa,ntwn auvtw/|\ 41 euri,skei ou-toj prw/ton to.n avdelfo.n to.n i;dion Si,mwna kai. le,gei auvtw/| Eurh,kamen to.n Messi,an o[ evstin meqermhneuo,menon Cristo,j\ 42 h;gagen auvto.n pro.j to.n VIhsou/n evmble,yaj auvtw/| o VIhsou/j ei=pen Su. ei= Si,mwn o uio.j VIwa,nnou su. klhqh,sh| Khfa/j o] ermhneu,etai Pe,troj
 
Dan,
See if this looks ok to you – the accents and diacriticals may need to be adjusted, I just guessed quickly to get a sample going, I’m too tired to think much…
I had an old sed script which was close enough that I just needed to change the values for the transliteration.
You can see the choices for composing (combining) characters here:
unicode.org/charts/PDF/U0300.pdf

1:1 ʼΕν ἀρχῃ̃ η̃̓ν ὁ λόγος καὶ ὁ λόγος η̃̓ν πρὸς τὸν θεόν καὶ θεὸς η̃̓ν ὁ λόγος
2 ου̃̔τος η̃̓ν ἐν ἀρχῃ̃ πρὸς τὸν θεόν
3 πάντα δι αὐτουͅ ἐγένετο καὶ χωρὶς αὐτουͅ ἐγένετο οὐδὲ έ̔ν ὸ̔ γέγονεν
4 ἐν αὐτῳ̃ ζωὴ η̃̓ν καὶ ἡ ζωὴ η̃̓ν τὸ φῳς τῳν ἀνθρώπων̀

This is the raw UTF8 (NFD mode) which regular expressions can work on from vim very nicely, but which generally look bad on web browsers and text only editors.
eg the accents will be a little to the right or left of where they ought to be because the web browser authors were too lazy to do it right – and therefore demand all unicode be in a pre-composed format (eg: characters which are not separate from their diacriticals…)

UTF8 can be “normalized” to a mode known as NFC (Normal form composed) where the accents are part of the letters themselves – which is not so good for doing analysis/searches by computer – but which will look right on Web pages, and simple text terminals – and if I understand you, that’s what you need.

I just load the text into graphical vim and it looks perfect when I choose a good font.

There are programs available to do normalization, but the four I found last night were all broken – and I am not functional enough to fix them right now. eg: It might be a month at the rate I am able to work…
If you know java, I believe in a version J2SE, version 6 or something? there is a class called Normalize which does exactly what you need – so if you poke around you might be able to write a 10 line program which cleans the text up.

Let me know what you think – and if you spot any gross errors, eg: I translated a character wrong, let me know and I’ll adjust the script.

–Andrew.
 
Dan,
See if this looks ok to you – the accents and diacriticals may need to be adjusted, I just guessed quickly to get a sample going, I’m too tired to think much…
I had an old sed script which was close enough that I just needed to change the values for the transliteration.
You can see the choices for composing (combining) characters here:
unicode.org/charts/PDF/U0300.pdf

1:1 ʼΕν ἀρχῃ̃ η̃̓ν ὁ λόγος καὶ ὁ λόγος η̃̓ν πρὸς τὸν θεόν καὶ θεὸς η̃̓ν ὁ λόγος
2 ου̃̔τος η̃̓ν ἐν ἀρχῃ̃ πρὸς τὸν θεόν
3 πάντα δι αὐτουͅ ἐγένετο καὶ χωρὶς αὐτουͅ ἐγένετο οὐδὲ έ̔ν ὸ̔ γέγονεν
4 ἐν αὐτῳ̃ ζωὴ η̃̓ν καὶ ἡ ζωὴ η̃̓ν τὸ φῳς τῳν ἀνθρώπων̀

This is the raw UTF8 (NFD mode) which regular expressions can work on from vim very nicely, but which generally look bad on web browsers and text only editors.
eg the accents will be a little to the right or left of where they ought to be because the web browser authors were too lazy to do it right – and therefore demand all unicode be in a pre-composed format (eg: characters which are not separate from their diacriticals…)

UTF8 can be “normalized” to a mode known as NFC (Normal form composed) where the accents are part of the letters themselves – which is not so good for doing analysis/searches by computer – but which will look right on Web pages, and simple text terminals – and if I understand you, that’s what you need.

I just load the text into graphical vim and it looks perfect when I choose a good font.

There are programs available to do normalization, but the four I found last night were all broken – and I am not functional enough to fix them right now. eg: It might be a month at the rate I am able to work…
If you know java, I believe in a version J2SE, version 6 or something? there is a class called Normalize which does exactly what you need – so if you poke around you might be able to write a 10 line program which cleans the text up.

Let me know what you think – and if you spot any gross errors, eg: I translated a character wrong, let me know and I’ll adjust the script.

–Andrew.
Thanks very much!

It looks very good. Only slight differences on some of the accents, such as on HN and hOUTOS but that may just be the way the font scales. Perhaps that is what you mean by normalization. I don’t currently have a DEV box for Java. What is the scripting language you are using?
 
Thanks very much!

It looks very good. Only slight differences on some of the accents, such as on HN and hOUTOS but that may just be the way the font scales. Perhaps that is what you mean by normalization. I don’t currently have a DEV box for Java. What is the scripting language you are using?
I am just using sed – stream editor. The standard command line utility for unix – it shows up on all flavors Sun, Linux, Caldera, etc. Some versions of sed can’t do escape sequences such as "
" for newline, etc – but you mentioned Linux and that has Gnu-sed which is bullet-proof.

Here it is wrapped in a shell script, chmod +x sedscript to make it executable.
Code:
#!/bin/bash
cat $1 | sed -e "y%abcdefghijklmnopqrstu!wxyz%αβχδεφγηιςκλμνοπθρστυ!ωξ!ζ%" -e "y%ABCDEFGHIJKLMNOPQRSTU!WXYZ%ΑΒΧΔΕΦΓΗΙΣΚΛΜΝΟΠΘΡΣΤΥ!ΩΞ!Ζ%" -e "s%\]%.%g" -e "s%/|%|/%g" -e "s%V%\xca\xbc%g" -e "s%v%\xcc\x93%g"  -e "s%,%\xcc\x81%g" -e "s%\`%\xcc\x94%g" -e "s%-%\xcc\x83\xcc\x94%g" -e "s%=%\xcc\x83\xcc\x93%g" -e "s%.]%\xcc\x80%g" -e "s%/]%\xcd\x85%g" -e "s%|]%\xcc\x83%g" -e "s%]%\xcc\x81\xcc\x94%g" -e "s%]]%\xcc\x80\xcc\x94%g" -e "s%;%\xcc\x93\xcc\x80%g" -e "s%:space:]][0-9]%
&%g"   > $2
It looks like I might have a loose iota subscript being placed under the omegas – I am not sure what the diacritics are supposed to be anyway, as I strip them except for rough breathing and the large tilde for my purposes – so take a good look.
Yes, Normalization to the NFC form will fix the font issues…

Unfortunately, the easiest to do normalization on is Java.
There is a perl script, but something was out of date and it failed to load the conversion tables properly…

search.cpan.org/~sadahiro/Unicode-Normalize-1.03/Normalize.pm

I also came across a php version, but it is just the class library and it was too convoluted for me to figure out last night. You would have to write a driver program to call the class subroutines … and the straight forward load file, dump through normalizer, and dump output program I wrote puked on several directory dependencies… etc. So I gave up and went to bed…

If you want to do some digging and working on these programs, I would be happy to give you pointers about what I know. I know how these things work – I just am unable to concentrate sufficiently to bring all the pieces together. Medical issues…

–Andrew.
 
Status
Not open for further replies.
Back
Top