And how are you going to verify what the poster meant in their mother tongue if you don’t know the language as you have already admitted? Just curious…
Yeah. :-p
What I do is run it through a couple of different translators, which has worked so far. Worst case, I’d have to ask for help.
Exactly as @ilfelice pointed out: if you don’t speak my native language, how will you moderate the source text? You will copy it and paste it into Google Translate or DeepL (which are neural networks/LLMs).
So the process is: I translate my thoughts using an AI, and then you use another AI to translate it back to verify if my AI made a mistake? This is a bureaucratic loop created just to justify an unworkable anti-AI policy.
Let’s treat each other like adults. I will continue to post in English, and if my phrasing ever sounds ambiguous or rude, just ask me to clarify. We don’t need digital “hide boxes” or mandatory bilingual declarations for that.
Do you now see the contradiction (and inviability) of what you are trying to enforce? ![]()
As a non-native English speaker, I disagree with this. I don’t think it is a good idea to rely on translation tools if you have some (even if limited) skills in English. There is always some nuance embedded in sentences you write directly that may not be carried through by translation tools. To prove this, I’ll post what I’d write in my native language and paste the Google translate below.
What I'd write in my native language
Mình cũng không phải là người bản ngữ, nhưng mình không nghĩ vậy. Mình nghĩ không nên dựa dẫm vào phần mềm dịch khi đã biết một chút tiếng Anh. Thường thì nếu mình tự viết thì sẽ có một số điều mà dịch máy không thể diễn đạt hết.
Google Translate
I’m not a native speaker either, but I don’t think so. I believe you shouldn’t rely on translation software when you already know a little English. Often, if you write something yourself, there will be some things that machine translation can’t fully convey.
That being said, I don’t agree with the suggested AI ban or mandatory disclaimer either. I think it is the individual’s responsibility for whatever content appearing under their name. It may be translated, generated by AI, or written by an entirely different person than the main account holder, but the account owner holds the ultimate decision on what to post and should therefore be accountable.
No, I see no contradiction.
A person posting their source text commits them to that text as what they meant to say. If the translated version that they post is wildly different from what multiple reverse translators come up with, it’s pretty obvious that any offensive parts in it aren’t translation errors.
Is that necessary for what we are proposing? Wouldn’t it be enough to ask the developers to provide a .hpkg package file for their software? A simple repo like others have (BeSly, FatElk etc..) I would just stipulate that they have open source available online on Github Gitlab etc..
Hmmm… Strange way of thinking. On the one hand we are being asked not to use LLM to translate our posts, and on the other you want to check the veracity of our posts in the original language using what is most likely an LLM based translation service. That seems pretty much the definition of contradictory, but I guess we can agree to disagree. ![]()
It also seems like a lot of extra burden for the moderators, which really sounds impractical (and probably not sustainable in the long term) for a small community like Haiku that struggles to find volunteers. But, what do I know. ![]()
That probably would work, yeah; you would be depending on the developer to certify that the package matched the published source, but if you’re ok with that then there shouldn’t be any technical problems.
Another difference is that LLM have no connection to real world experience at all, all they can do it perceive (sorry for anthropomorphism here, but analogy with image recognition is useful) word patterns.
The way these things process text is very unlike how humans think.
Having the “source language” together with the translation isn’t just useful for moderators. I wouldn’t expect moderators to examine each and every source-text (or get help from a speaker of that language). It may only become helpful if some controversy arises and we can have a look if the translation contributed.
The source-text may also be useful for other speakers of that language. They may point out that some misunderstanding is on account of a bad translation. Doesn’t even have to be a big scandal, just a misunderstanding that leads everyone wanting to help into the wrong direction. ![]()
I think non-LLM-based translation tools are fine, and I am not trying to ban those. The LLM-based ones (Google Translate currently has an option for Gemini vs. “Classic” translation) are clearly much worse at translating what one actually said, and instead spit out some vague approximation of it. To be clear, I also don’t think “neural network” algorithms are inherently a problem either. Machine translation, when it’s done “mechanistically”, I think is fine. This is not trying to interpret speech, but render it relatively literally (or sometimes a little less than literally, if it has “idiom dictionaries” and can deal with things like that) into another language.
If the machine translation actually tries to convert meaning and not just words across languages, not using dictionaries (or neural networks trained to act like more advanced dictionaries), this is where I think there is a problem.
Also: “purity” isn’t a word I used, nor any close synonyms I don’t think. So where are you getting that from…? Human speech is often “impure” by many metrics. But it still represents people’s actual thoughts, which are themselves often unrefined.
I think the idea here is that there are other speakers of many other languages besides English on this forum, and if one posts both the original language as well as the machine translation (as Andrea has started doing), then those other speakers of your language can pick up on details the machine translation may have not carried over so well, and we can all communicate better. I am not really fluent in any language besides English, but I know some Spanish and Italian, enough that I probably will be able to glean details now and again from those that any machine translation will lose.
So the point isn’t to create an “English-only club”, rather it’s to make it easier to use non-English languages for communication.
I actually know at least one person in real life who I trust who speaks Russian, and if actually necessary I could consult them. I suspect there are others on this forum and elsewhere who know enough that we could also consult, too.
Besides that, we can ignore the machine translation software, and do word-by-word lookups in lexicons and dictionaries and see for ourselves what each word is noted to mean, and find example usages and infer connotations from that, same as anyone learning a second language can.
I am trying to describe the experience of thinking independent of whatever phenomena is ‘behind’ it. Careful examination of our mental phenomena and comparison with that of others, independent of however these things “work”, may reveal facts we would have a hard time getting at any other way.
The moment we start talking about “neural pathways” in relation to cognition, we are actually speaking somewhat unscientifically because neuroscience cannot fully explain human cognition at present. To state that “LLMs work similarly” is then, I think, an assumption that there is insufficient evidence to consider even a hypothesis.
I strongly disagree with this way of thinking because I strongly believe that all that Plato’s categories are just ancient fairy tales and make no sense today, it only exists as information in our brain. We have invented more strict theories like formal logic of predicates and math set theory that is a foundation of all modern science, technology and computing. I do not want Haiku to become a specific school of philosophy and reject everyone who have other views of our world and reality. Haiku is about making software, not about understanding truth about world inner workings.
World works with duck typing, not with ideal categories.
How does this actually go beyond Aristotelian logic? Sure, it may be easier to work with in mathematical terms, but in what way is it actually beyond the classic system of syllogisms?
Except Godel’s incompleteness theorem means it’s not really a basis for mathematics in the way Russel et al. were trying to make it be in the early 20th century.
And besides, “set theory” may be useful in many ways, but as an actual “foundation” of mathematics, well, it’s not. Wittgenstein has a pretty good explanation of how it can’t be (he’s here speaking about the Principia mathematica, Russel and Whitehead’s attempted formalization of all mathematics, but the exact same thing applies to ZFC set theory or any other attempted basis):
If I give you a calculation to do, you say that you will do it by Principia. But what if I do it in the ordinary way and get a different result? How do we decide which calculation is correct? … And we might trust one thing although the other disagreed with it–we might even have to say, for example, that Russell’s logic gives wrong results. We cannot say that arithmetic is based on Russell if Russell is based on arithmetic.
(Wittgenstein, “Lectures on the Foundations of Mathematics”.)
I can understand that. Until now, it was quite possible for us all to disagree about all the philosophical matters that did not affect the direction of Haiku’s development without it causing strife and disagreements. But now, well, we have encountered something which some of us, for philosophical reasons (among others), want nothing to do with it, not just for ourselves but in the projects and communities we are in and work on.
We have attempted to find technical reasons (e.g. unreliability, copyright, etc.), in the hope that we could agree about those and leave the philosophical disagreements to the side once more. But it seems so far we have failed to do that. And failing that, well, we are bringing out all our philosophical reasons, to see if those would convince; and even if they won’t find agreement, that they will at least make clear that we do have strong reasons for what we are saying and doing. And barring something that convinces us otherwise, we aren’t likely to change our minds about what we are going to accept.
But ducked-typed languages such as JavaScript still have actual, definite types underneath the “duck” behavior. It just converts between them more easily than most other languages.
There seems to be multiple LLM threads here and i am confused where to put this. Well, if it is in wrong place.. move it freely.
I use LLM nowadays with my own hardware, 20b models run reasonable speed and i have different models for different purposes. I use it as help to solve different problems in my projects and as a “personal tutor” for learning some new things. Also translating text for some projects..
I dont use LLM anywhere else. I have Nokia 3210 4G as my main phone, I dont have Windows anymore anywhere (only Linux and Haiku, ohh and 14 years old iMac 27” saved from trash and fixed it), my TV is 16 years old and very stupid (couple new capasitors replaced once and still going strong), my cars are 42 years old average, now i drive with 40 years old Hiace as daily driver. I like simple (stupid) analog old school things and also like to fix old things back to life.
I hate companies which tries to shovel AI everywhere and are forcing people to use it and scrape all data from internet and make websites servers choke because of it.
But, about local LLMs, i like those as a personal tool. It gives me possibilities to do thing which i probably cannot do without it, but i also try to learn things at the same time.
My local LLM helped me to get my gaming rig/hardware to work in games with Linux, it helped to built scripts and python apps to get some things to work, which i didnt never got working in windows. All games and controllers (also simulators) work now perfectly in my Linux machine. It also helped to solve some problems with my t460s Haiku laptop, helped me to make few little apps to it also.
Yeah, this LLM hype divides people totally. I am fine with that. Everybody doesnt have to like it, but i also understand its possibilities because it works for me with stuff i like to do. It can be used good and bad.. like everything else.
I also know couple of LLM hardcore fanatics which are using local “LLM agents” with their monstrous computers which they own. They gives exact intructions and plans to the agents and are running multiple pipelines through day and night and later they review the results while drinking beer or whatever.. i am fine with that also. If they find that useful so be it.
One guy runs his local LLM rig with self trained models and solar energy and do some medical research stuff with it..
Same for many of us with many aspects of the world. For example, I have never been to the US and have no “real world experience” of the place. All I can do is read streams of text and images from the news, which invokes certain thoughts about that country based on my experience in other countries.
The same applies to all of us about the past - should every historical analysis by a living human being be disregarded then?
I agree with this. The full human experience is much more complex than this.
But how is it unscientific? Aren’t there lots of studies where people are presented with certain triggers or asked to do certain tasks, then consistent brain scan patterns are revealed?
It is kinda like how compilers can’t be explained by “string parsing”. But in some cases, a bunch of regex patterns can go quite a long way. They may fail for all but the very simple inputs, sure, but we can’t just automatically disregard the output because they do not follow the established practices of modern compilers.
No, it does not. Machine translation is OK, but using a chatbot LLM is not.
Nobody asked you to see what kind of tech your machine translation uses, the policy is very clearly talking about LLMs directly, (though, you should also not use a translation service that “just uses gpt” or the like) i.e the usecase of copy pasting your text into chatgpt and writing “make this english”, the text will be heavily distorted and we now have to clean up this mess and suffer it, this is why this is forbidden.
Translation services (like vertalen.nl for example) do not have this problem, they don’t significantly alter speach, and they aren’t capable of creating a giant wall of text in the sense of “just write a response to this!”
I think what you don’t notice here, is that asking “noob” questions is actually a form of contribution. It highlights blind spots in the documentation, that we then try to find. If everyone instead turns to LLMs ask to ask these questions, the documentation will be allowed to be more and more incomplete, and the information not directly accessible, making us more dependant on LLMs for onboarding new contributors.
This is similar to projects that don’t have any online documentation and instead a “join our discord server and ask” link.
So, in that case, the rule would be: if you had to ask an LLM, consider why, and see if the missing information could be fixed in some other way.
It also opens extra possibilities for misunderstandings (intentional or not). Israel government is known to do that in their press releases: the english and hebrew versions are slightly different, in ways that you can assign to the difficulty of translation, but that seem intentional.
I would prefer to have only one version of the text. And if people choose to use a tool that misrepresents their thoughts, it’s their problem, not ours as readers? I think we can let each of us decide if we are more comfortable writing in english directly, or if using translation, what tool to use. Even if you don’t speak english, you can test these tools by doing round-trips of the translation.
The only thing I eould rather not have to read is people who prompted an LLM to generate a forum post (rather than just doing a translation). But we already banned that.
This confusion is on me for not have written clearly; my apologies.
Sure, I understand the logic, but how would you know if a post was translated by LLM or not? Realistically, this is not enforceable, unless you police every post, and interrogate every time you become suspicious. More burden for the poor moderators… ![]()
I know three languages (Spanish, English and Japanese) pretty well. My experience is completely the opposite of yours: legacy (pre-LLM) were pretty primitive and not always accurate, while the current LLM-based translators (which I use quite often) are way better. What languages is your experience with LLM-based translators based on? Just curious.