طيب لماذا النتيجة خاطئة في التنفيذ
تنفيذ البرنامج بشكل متوازي على معالجين بنفس الجهاز ارجو الافادة
تم تعديل هذه المشاركة بواسطة aboazoz2004 في 27 مايو 2010 في 15:05
Khaled.Alshaya كتب:معك حق, أنا لا أعلم عن OpenMP, و لكن هناك Syncronization Problem في الكود الذي وضعته, لأن جميع الـ Threads التي سوف تظهر تتشارك في شيء, ماهو هذا الشي؟!
أنا أسميها Multiproblems بدلاً من Mutliprocessing
1+ لأول شخص يخبرنا :)
تحياتي...
اخي الكريم اذا انت لا تعلم بال OpenMP فكيف قلت ان هنالك Syncronization Problem ؟؟؟ ا يوجد اي خطاء
في ال OpenMP وضيفه ال #pragma omp for ordered هي توزيع العمل بالتساوي على ال processors .
انا سوف اعدل حجم المصفوفه واطبعلك الناتج على جهاز core 2 Duo
#include<iostream>
#include <stdio.h>
#include <conio.h>
#include <omp.h>
using namespace std;
int main()
{
int Array[100];
for (int i=0;i<100;i++)
Array=i;
#pragma omp parallel
{
bool IsPrime;
#pragma omp for ordered
for (int i=0;i<100;i++)
{
IsPrime=true;
for(int j=2;j<=Array/2;j++)
if(Array%j==0)
{
IsPrime=false;
break;
}
if(IsPrime)
printf("%d\n",Array);
};
}
getch();
}شاهد البرنامج وصوره ال output
اخي اذا انت تعرف فعلاً اي الخطاء الذي تتحدث عنه وضحه اين هو ؟؟؟ ولا فان النقاش من غير فائده اذا لم تحدد اين مشكلتك ؟
الـ std output عبارة عن resource مشترك, هو ليس thread safe في كل من C و ++C, لذلك يجب أن تقوم بعمل Synchronization لكي يصبح الكود صحيح,
معنى أنه عمل على Implementation معين على نظام تشغيل معين حيث الـ std io معرفة على أنها thread safe, لا يعني أن الكود صحيح.
جيد
طيب كيف اعرف عدد المعالجات على جهاز
انا شغال على core2 duo
لكن اريد ان اعرف كم معالج عندي على جهازي ليس بالمحاكاة انما فعليا
لماذا يوجد مشكلة اذا زاد حجم المصفوفة
اقتباسجيدطيب كيف اعرف عدد المعالجات على جهاز
انا شغال على core2 duo
لكن اريد ان اعرف كم معالج عندي على جهازي ليس بالمحاكاة انما فعليا
الطرق التي أعرفها تخبرك كم عدد الـ Computational Units الموجودة في الجهاز الذي يعمل عليه البرنامج. بمعنى آخر, كم عدد الـ Cores و الـ Hyperthreading Units. عموماً, انظر لهذا الموضوع لتعرف الطريقة لأي نظام تريده:
http://stackoverflow.com/questions/150355/programmatically-find-the-number-of-cores-on-a-machine
بالنسبة لي, أرى أن العمل قد تم إنجازه من قبل باستخدام boost:
std::size_t number_of_computational_units = boost::thread::hardware_concurrency();
الآن هل تعتقد أنك لو أنشأت 4 threads على جهاز يحمل 4 cores سوف يتم تنفيذهم كـ Parallel Threads؟ ليس هناك شرط على الإطلاق, هذا يرجع لنظام التشغيل و كمية البرامج التي تعمل على الجهاز, و في الغالب, سوف يكون تشغيل البرنامج Sequentially أسرع, لأن هناك عمليات IO تحصل على الدوام و كل الـ Threads تستخدمها في نفس الوقت.
هل نستطيع القول للكود السابق انه تم تنفيذه بشكل متوازي على اكثر من معالج
السلام عليكم ...
أخيراً وجدت بعض الوقت :)
قمت بتنفيذ طلبك بحذافيره, و هو معرفة عدد الـ Computational Units الموجودة على الجهاز, و من ثم قمت بإنشاء threads بنفس العدد, مو لكن تركت الـ main thread غير مشغولة لكي يبقى البرنامج منفصلاً عن الكود في حالة استخدامه كقطعة في نظام أكبر. البعض و أنا منهم يفضل عدم التعامل مع الـ Threads يدوياً, و من الأفضل أن تستخدم High Level Multiprocessing و الـ Threads تعتبر الـ Low Level Side في هذا العالم, و لكن للفائدة لا أكثر. استخدمت boost, و المثال المرفق يعمل من سطر الأوامر, يمكنك بالطبع تعديله ليناسب احتياجاتك. بالنسبة للتعليقات, هي كثيرة و لكن هذه هي طريقتي في كتابة الكود :)
#include <cstddef>
#include <iostream>
#include <algorithm>
#include <exception>
#include <boost/thread.hpp>
#include <boost/lexical_cast.hpp>
/* Welcome to the real world of "Multiproblems"!
The standard streams are not thread-safe!
*/
std::vector<int> numbers;
boost::mutex printing_guard;
boost::thread_group threads;
/* This is a very bad way of testing for primes,
I think you have to "rethink" the way as a whole,
because this is a very hard problem actually(prime test).
Look at Sieve of Eratosthenes if you have a reasonable range to test */
bool is_prime(int number)
{
// primes start at 2
if(number <= 1) return false;
for(int possible_factor = 2; possible_factor < number; ++possible_factor)
{
if(number % possible_factor == 0) // number is not prime number
return false;
}
return true;
}
void print_if_prime(int number)
{
if(is_prime(number))
{
printing_guard.lock();
std::cout << number << std::endl;
printing_guard.unlock();
}
}
void is_prime_range(const std::vector<int>& numbers, std::size_t beg, std::size_t end)
{
std::for_each(numbers.begin()+beg, numbers.begin()+end, print_if_prime);
}
int main(int argc, char* argv[])
{
/* get the numbers from the command line. For example:
$ primes 1 2 3 4 5 6 7 8 9 10
output:
2
3
5
7
*/
numbers.resize(argc-1);
try
{
std::transform(argv+1, argv+argc, numbers.begin(),
boost::lexical_cast<int, char*>);
}
catch(const std::exception&)
{
std::cout << "Error: You have entered invalid numbers.";
}
/* After we knew how many computational units we have
on a specific machine, we create an equal number of threads,
and we divide the work equally between them. If the work can't
be divided equally, we assign the remaining data to a thread we call man thread.
This case happens, when the "number" of "numbers" has a reminder modulus the number of threads.
*/
/* Boost is your friend here, don't reinvent the wheel :)
boost::thread::hardware_concurrency() returns the number
of computational units on the current machine.
*/
/* There is a main thread also! keep it responsive at least,
we are working on three other CPUs on my machine. */
std::size_t number_of_threads = boost::thread::hardware_concurrency() - 1;
std::size_t numbers_count = numbers.size();
std::size_t block_size = numbers_count / number_of_threads;
std::size_t block = 0, cpu = 0;
if(block_size) // if there is work to divide
{
// keep one thread for later usage, because we want
// to assign the rest of the data to the last thread,
// in the case that work doesn't divide equally between threads
for( ; cpu < number_of_threads-1; ++cpu, ++block)
{
threads.add_thread( new boost::thread(is_prime_range, numbers, block*block_size, (block+1)*block_size) );
}
}
threads.add_thread( new boost::thread(is_prime_range, numbers, block*block_size, numbers.size()) );
threads.join_all();
}طريقة استخدامه بسيطة, استدع البرنامج من سطر الأوامر مع الأعداد التي تريد التحقق منها(للتجربة), و سيطبع الأعداد الأولية:
$ primes 1 2 3 4 5 6 7 8 9 10 2 3 5 7 ---------------------------- Note That primes don't have to be ordered, because threads are working in parallel. This is just a hypothetical output.
بالطبع يمكنك وضع الـ primes في vector و من ثم نقوم بترتيبهم لكي نطبعهم بشكل مرتب في النهاية:
#include <cstddef>
#include <iostream>
#include <algorithm>
#include <iterator>
#include <exception>
#include <boost/thread.hpp>
#include <boost/lexical_cast.hpp>
/* Welcome to the real world of "Multiproblems"!
The standard streams are not thread-safe!
*/
std::vector<int> numbers;
std::vector<int> primes;
boost::mutex vector_guard;
boost::thread_group threads;
/* This is a very bad way of testing for primes,
I think you have to "rethink" the way as a whole,
because this is a very hard problem actually(prime test).
Look at Sieve of Eratosthenes if you have a reasonable range to test */
bool is_prime(int number)
{
// primes start at 2
if(number <= 1) return false;
for(int possible_factor = 2; possible_factor < number; ++possible_factor)
{
if(number % possible_factor == 0) // number is not prime number
return false;
}
return true;
}
void keep_if_prime(int number)
{
if(is_prime(number))
{
vector_guard.lock();
primes.push_back(number);
vector_guard.unlock();
}
}
void is_prime_range(const std::vector<int>& numbers, std::size_t beg, std::size_t end)
{
std::for_each(numbers.begin()+beg, numbers.begin()+end, keep_if_prime);
}
int main(int argc, char* argv[])
{
/* get the numbers from the command line. For example:
$ primes 1 2 3 4 5 6 7 8 9 10
output:
2
3
5
7
*/
numbers.resize(argc-1);
try
{
std::transform(argv+1, argv+argc, numbers.begin(),
boost::lexical_cast<int, char*>);
}
catch(const std::exception&)
{
std::cout << "Error: You have entered invalid numbers.";
}
/* After we knew how many computational units we have
on a specific machine, we create an equal number of threads,
and we divide the work equally between them. If the work can't
be divided equally, we assign the remaining data to a thread we call man thread.
This case happens, when the "number" of "numbers" has a reminder modulus the number of threads.
*/
/* Boost is your friend here, don't reinvent the wheel :)
boost::thread::hardware_concurrency() returns the number
of computational units on the current machine.
*/
/* There is a main thread also! keep it responsive at least,
we are working on three other CPUs on my machine. */
std::size_t number_of_threads = boost::thread::hardware_concurrency() - 1;
std::size_t numbers_count = numbers.size();
std::size_t block_size = numbers_count / number_of_threads;
std::size_t block = 0, cpu = 0;
if(block_size) // if there is work to divide
{
// keep one thread for later usage, because we want
// to assign the rest of the data to the last thread,
// in the case that work doesn't divide equally between threads
for( ; cpu < number_of_threads-1; ++cpu, ++block)
{
threads.add_thread( new boost::thread(is_prime_range, numbers, block*block_size, (block+1)*block_size) );
}
}
threads.add_thread( new boost::thread(is_prime_range, numbers, block*block_size, numbers.size()) );
threads.join_all();
std::sort(primes.begin(), primes.end());
std::copy( primes.begin(), primes.end(),
std::ostream_iterator<int>(std::cout, "\n"));
}هل هذا هو طلبك؟ لا تنسى أنك تحتاج إلى أكثر من core لكي يعمل البرنامج(أعني هنا multiple threads) !
تم تعديل هذه المشاركة بواسطة Khaled.Alshaya في 27 مايو 2010 في 22:08
اشكرك لكن لم يتم الرد بخصوص النتيجة لماذا تظهر مختلفة عن التنفيذ العادي اقصد النتيجة خاطئة حتى لو لم تكن بالترتيب
حتى الان لم يتم الرد
اقتباسحتى الان لم يتم الرد
كيف لم يتم الرد؟!
اقتباسالآن هل تعتقد أنك لو أنشأت 4 threads على جهاز يحمل 4 cores سوف يتم تنفيذهم كـ Parallel Threads؟ ليس هناك شرط على الإطلاق, هذا يرجع لنظام التشغيل و كمية البرامج التي تعمل على الجهاز, و في الغالب, سوف يكون تشغيل البرنامج Sequentially أسرع, لأن هناك عمليات IO تحصل على الدوام و كل الـ Threads تستخدمها في نفس الوقت.
ما تريد فعله, لم أرى أحداً يقوم به. لديك فكرة خاطئة عن الـ Multiprocessing. إذا كنت تريد فعل ما تريده, عليك أخذ نظام تشغيل كـ Linux, و من ثم تقوم بالتعديل على الـ Dispatcher, و تكتب ما تريد هناك. هذه هي الطريقة الوحيدة المضمونة. و لكن لماذا كل هذا العناء, و ما أدراك أن الـ Net Performance سوف يكون أكبر؟ هذه هي الفكرة, نظام التشغيل يقوم بعمل تنظيم للـ Performance لجعل العملية أكثر كفاءة مما لو قام كل مبرمج بالتحكم يدوياً لأن هناك أكثر من برنامج يعمل في نفس الوقت بكل تأكيد. هذه هي أنظمة التشغيل منذ السبعينيات إلا في حالة الـ Embedded Systems و الـ Super Computers!
تم تعديل هذه المشاركة بواسطة Khaled.Alshaya في 5 يونيو 2010 في 23:32

